{
  "id": 13731405,
  "title": "These execs think voice AI hasn’t reached its ChatGPT moment yet",
  "url": "https://urgent.news/2026/10/11/these-execs-think-voice-ai-hasnt-reached-its-chatgpt-moment-yet",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T14:00:00.000Z",
  "source": {
    "name": "TechCrunch",
    "slug": "techcrunch",
    "url": "https://techcrunch.com/2026/10/11/these-execs-think-voice-ai-hasnt-reached-its-chatgpt-moment-yet/"
  },
  "original_language": "en",
  "account": "Voice AI, a burgeoning interface technology, has recently garnered significant investor interest, with billions of dollars being directed towards various startups working on diverse applications such as model creation, enterprise customer service, and meeting note-taking. The industry is awash with new models and tools claiming to emulate human conversation. However, despite the advent of full-duplex models that allow for simultaneous speaking and listening, PolyAI's Chief Technology Officer Shawn Wen contends that voice AI has yet to reach the \"ChatGPT moment.\"\n\nWen's primary concern lies in the speed of reasoning within these AI models. He believes that the next hurdle is to make these systems capable of swiftly fetching answers, thereby creating a more natural conversational experience. Furthermore, Wen stresses the need for enterprise voice AI agents to sound less robotic and instill user confidence in their ability to solve problems.\n\nThe conversation turns to the realm of meeting notetaking, where Otter, a meeting note-taking platform, is making strides. CMO Alex Gay highlights the importance of speaker identification, intent capture, and the incorporation of organizational knowledge into AI notetakers. He emphasizes that accurately reflecting the emotive expressions and dynamics of human conversations is crucial for the acceptance of digital twins in meetings.\n\nHowever, despite advancements in voice AI, AI assistants continue to stumble, often failing to grasp user queries or failing to accurately transcribe meetings. Wen identifies Automatic Speech Recognition (ASR) models as the culprit, noting their propensity to overlook critical keywords, which in turn affects the overall context of the conversation. Similarly, Otter's Gay underscores the necessity for continuous improvement in ASR models, citing the inherent flaws that arise from inaccurate transcriptions and the cascading effects on downstream actions.\n\nMoreover, the transparency of voice AI tools is a point of contention. Both PolyAI's Wen and Otter's Gay advocate for clear declarations of AI involvement in conversations, including the notification of participants when a meeting is being recorded. This push for transparency aims to foster trust, especially in enterprise settings where AI agents engage in calls.",
  "summary": "Voice AI's often misses important points for its context layer, and causes the whole pipeline to break",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}