Localization is the missing test for voice AI
Gartner predicts that by 2028, 70% of customer journeys will begin with a conversational AI interface.
The voice AI industry is experiencing rapid expansion, with projections indicating it could reach $47.5 billion by 2034, up from $2.4 billion in 2024. Gartner predicts that by 2028, 70% of customer service interactions will begin with a conversational AI interface. For multinational corporations, the appeal of voice AI is clear - it can provide consistent, scalable customer support across diverse markets, languages, and time zones.
However, there is a potential issue that voice AI may encounter, similar to the challenges faced by outsourced customer service centers. Differences in language and culture can create disconnects and frustration between the brand and the customer. As voice AI continues to advance, language may become less of a barrier, but merely speaking the same language does not equate to truly understanding someone.
Customers bring accents, habits, and cultural expectations to every interaction, and they expect support to reflect their reality. Localization should be integral to assessing voice AI performance, rather than being considered a mere translation process conducted at the end of development. Accents, dialects, terminology, and cultural norms can vary significantly even when speaking the same language.
For example, in Scotland, a customer might say "aye" instead of "yes," refer to something small as "wee," or talk about "getting the messages" when they mean going shopping. The language may be English, but understanding the interaction requires familiarity with how that language is used locally. Cultural differences also affect expectations around directness or politeness.
The way a customer is addressed can be crucial in some cultures, while in others, addressing someone by their first name is the norm. These factors are not merely cosmetic details; they impact an AI agent's ability to comprehend intent and respond appropriately, ultimately determining whether customers trust the experience provided.
Voice AI has made significant strides, as evidenced by companies like ElevenLabs, which has developed AI-generated voices that reproduce speech pacing, intonation, and emotion, making synthetic voices sound increasingly human. However, as these technologies become more convincing, the line between human and machine becomes increasingly blurred.
This shift in perception extends beyond customer service, with even actors like Matthew McConaughey embracing AI voice technology to replicate their performances without altering their natural voice. Nevertheless, not all AI voices are uniformly advanced, and many people can still identify underdeveloped AI voices. Moreover, humans are adept at detecting slight imperfections, which can disrupt the illusion of a natural conversation.
This challenge is further exacerbated by local context. A voice can sound remarkably human while still feeling culturally out of place. As voice AI continues to approach human-like speech, these incongruences become more noticeable. Global brands cannot rely solely on a one-size-fits-all approach for voice AI. While a frontier model can provide the foundation, its performance must be tested against local data and genuine customer interactions to determine what constitutes success in terms of empathy, terminology, and intonation.
These aspects cannot be inferred by a frontier model; they must be learned on a per-market basis, brand by brand. Successful voice AI models will have a continuous feedback loop that monitors interactions within individual markets, identifies instances of misunderstanding or friction, refines the experience, and tests the model again.
Local teams should actively participate in this process, bringing market expertise into development rather than merely receiving the "finished product." Voice AI cannot be treated as a static product; it requires ongoing refinement and adaptation. As language and slang evolve, customer expectations shift, and new trends introduce unfamiliar expressions and behaviors, the original training data may no longer be sufficient.
Neglecting localization can have serious consequences for a brand's reputation. The recent experience of Derby City Council with its AI assistant, Darcie, serves as a cautionary tale. Despite being upgraded to support nine additional languages, Darcie struggled to understand a presenter with a strong Derbyshire accent who used local expressions such as "mardy" and "duck."
This example underscores the importance of localization in ensuring voice AI meets the needs and expectations of diverse customer bases.
Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.