← บทความทั้งหมด
เปรียบเทียบ

WhatsApp Calling API มีบันทึกเสียงและถอดเสียง: แยก chat intake ออกจาก call intelligence

ใน changelog ของ WhatsApp Business Platform วันที่ 30 มิถุนายน 2026 Meta ส่งสัญญาณสำคัญให้ทีมซัพพอร์ต: Cloud API มีคู่มือ call recording และ call transcription ภายใต้ WhatsApp Business Calling API แล้ว พื้นผิว official calling นี้ยังครอบคลุม user-initiated calls, business-initiated calls, SIP, calling webhooks และ calling pricing

ถ้า workflow ซัพพอร์ต WhatsApp ของคุณมีเสียง นี่คือข่าวดี การบันทึกเสียงและ transcript ทำให้บทสนทนากลายเป็นหลักฐานที่ค้นหาได้: ลูกค้าถามอะไร เจ้าหน้าที่รับปากอะไร และต้อง follow-up อะไรต่อ แต่ไม่ควรตีความว่า transcript ของสายโทรได้แก้ปัญหา customer context ทั้งหมดแล้ว

มันแก้ปัญหา artifact ของเสียง ไม่ใช่ปัญหา chat intake

Call intelligence ไม่ใช่ message intake

Transcript ของสายโทรเกิดขึ้นระหว่างหรือหลัง voice session ใช้ดีสำหรับ QA, summary, training และ follow-up ส่วน pipeline chat inbound มีหน้าที่อีกแบบ: รับข้อความลูกค้าให้เร็ว ตรวจลายเซ็น delivery เก็บ event แล้ว route ไปยัง agent, CRM, queue หรือ AI worker

สองชั้นนี้ควรมาเจอกันใน customer timeline เดียวกัน แต่ไม่ควรเป็น dependency เดียวกัน

คำถามWhatsApp Calling API recording/transcriptionSigned inbound chat webhook
สิ่งหลักสายโทรและ recording/transcriptEvent ของข้อความลูกค้า
เวลาเกิดระหว่างหรือหลังสายโทรทันทีที่ข้อความ inbound มาถึง
เหมาะกับQA, summary, compliance review, coaching, follow-up notesIntake, routing, deduplication, CRM logging, AI triage
ความเสี่ยงที่ต้องกันเสียบริบทหลังจบสายข้อความแรกหายก่อนเข้า queue
ขอบเขตช่องทางWhatsApp voice ผ่าน official calling surfaceWhatsApp, Telegram, LINE, TikTok, Zalo และ X ใน event stream มาตรฐานเดียว

ถ้าทีมของคุณรับ WhatsApp voice เป็นหลัก official Calling API อาจเป็นศูนย์กลางของ workflow ได้ แต่ในไทยและญี่ปุ่น LINE เป็นช่องทางหลักมาก ส่วนทีมภูมิภาคยังมักมี WhatsApp, Zalo, Telegram และ TikTok ด้วย ในกรณีนั้น transcript ของสายโทรเป็นเพียง artifact หนึ่งใน timeline ที่กว้างกว่า ไม่ใช่ทางเข้าหลักของระบบซัพพอร์ต

ขีดเส้นแบ่งก่อน

Official WhatsApp Calling API ควรอยู่ในชั้น voice ใช้เมื่อผู้ใช้ต้องการโทร เมื่อคุณต้องมี call button เมื่อ agent ต้อง route สาย หรือเมื่อ QA ต้องใช้ recording และ transcript

Message intake ควรอยู่ในชั้น event ชั้นนี้ควรเรียบง่ายและเข้มงวด:

  1. รับ inbound delivery
  2. ตรวจลายเซ็น
  3. เก็บ raw event ตาม ID
  4. Route ตาม provider, account, conversation, sender และ message type
  5. ให้ AI, CRM และ human queue ทำงานหลังจากเก็บแล้ว

เรื่องนี้สำคัญเพราะ chat เป็น asynchronous ลูกค้าอาจส่งว่า “หลังจากคุยโทรศัพท์ เวลาไปรับของเปลี่ยนไหม” แล้วออกไป ถ้าเส้นทาง inbound ยังรอ transcript ของสายโทร, CRM lookup หรือ AI summary ระบบอาจพลาดข้อความแรกที่ควรเก็บทันที

UnifyPort อยู่ตรงไหน

UnifyPort ไม่ได้แทนที่ WhatsApp Calling API อย่างเป็นทางการ ถ้าคุณต้องใช้ WhatsApp voice calling, call recording หรือ official call transcription ให้ใช้ชั้น calling ของ Meta สำหรับส่วนนั้น

บทบาทของ UnifyPort แคบกว่า: รับข้อความ inbound จากบัญชี messaging ทั่วไปผ่าน unofficial interface แล้วส่งเป็น webhook event ที่ลงลายเซ็นและ normalize แล้ว handler เดียวสามารถรับ WhatsApp, Telegram, LINE, TikTok, Zalo และ X ได้

สร้าง webhook endpoint, subscribe message.received และตั้ง signing_secret:

curl -X POST https://api.unifyport.ai/v1/webhook-endpoints \
  -H "X-Api-Key: <YOUR_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
  "url": "https://support.example.com/webhook",
  "status": "active",
  "subscribed_events": ["message.received"],
  "signing_secret": "<WEBHOOK_SIGNING_SECRET>"
}'

เมื่อมีข้อความ WhatsApp เข้ามา receiver จะได้ event มาตรฐาน:

{
  "id": "evt_20260710_01",
  "type": "message.received",
  "provider": "whatsapp",
  "account_id": "acc_support_whatsapp",
  "occurred_at": "2026-07-10T02:30:00Z",
  "data": {
    "conversation": { "id": "84901234567", "type": "user", "title": "Minh Tran" },
    "sender": { "id": "84901234567", "type": "user", "name": "Minh Tran" },
    "message": {
      "id": "wamid.HBgM20260710",
      "type": "text",
      "text": "Can someone confirm whether my pickup changed after the call?",
      "direction": "inbound",
      "sent_at": "2026-07-10T02:29:58Z"
    },
    "event": { "kind": "message_received" }
  }
}

ถ้า endpoint มี signing_secret delivery จะมี X-Device-Timestamp และ X-Device-Signature ลายเซ็นคือ hex HMAC-SHA256 ของ <X-Device-Timestamp>.<raw request body> ทำให้ intake service ตรวจแหล่งที่มาได้ก่อนเขียนลง storage หรือส่งต่อให้ AI

สถาปัตยกรรมควรเรียบง่าย:

WhatsApp / LINE / Zalo / Telegram / TikTok / X message
  -> UnifyPort message.received webhook
  -> HMAC-SHA256 signature verification
  -> Event store
  -> Routing, CRM lookup, AI triage, or human queue

WhatsApp voice call
  -> Official WhatsApp Calling API layer
  -> Recording / transcript / summary artifact
  -> Attach to the same customer timeline

ประเด็นคือ customer timeline ควรเริ่มที่ไหน ควรเริ่มจาก event inbound แรก แล้วค่อยแนบ recording, transcript และ summary เป็น artifact ที่เกี่ยวข้อง

ทีมเล็กควรตัดสินอย่างไร

สำหรับทีมซัพพอร์ต 2-10 คน คำถามไม่ใช่ “ควรใช้ call transcription ไหม” แต่คือ “ระบบไหนเป็นเจ้าของ record แรกของลูกค้า”

ถ้า WhatsApp voice เป็นช่องทางหลัก official calling stack อาจเป็นศูนย์กลางได้ Agent รับสาย recording และ transcript ผูกกับสายโทร workflow เป็น voice-first

ถ้า chat เป็นช่องทางหลัก ชั้น intake ไม่ควรรอ voice tooling ให้เก็บข้อความก่อน แล้วค่อย enrich:

  • ถ้าลูกค้าโทรมาภายหลัง ให้แนบ recording และ transcript กับ conversation เดียวกัน
  • ถ้า AI draft คำตอบ ให้เก็บ message.received เดิมเป็น source
  • ถ้า agent ตอบ ให้ส่งผ่านบัญชีที่เชื่อมต่อด้วย POST /v1/messages
  • ถ้า LINE, Zalo หรือ Telegram เข้า queue เดียวกัน ให้ route ด้วย provider แทนการสร้าง inbox ใหม่

อัปเดต recording และ transcription ของ Meta ทำให้ WhatsApp voice มีประโยชน์ขึ้นสำหรับงานซัพพอร์ต แต่ไม่ได้แทนที่ชั้น chat intake ที่สะอาด แยกสองชั้นนี้ออกจากกัน แล้วทีมจะใช้ทั้งสองอย่างได้โดยไม่ให้ชั้นหนึ่งกลายเป็นคอขวดของอีกชั้น