Urgent.News

What's breaking now, across thousands of outlets.

AI

What happens to language when machines start talking back?

A few months ago, a friend posted something that caught my eye on social media. She struggled to listen to what she termed "(literally) breathless AI voices for audiobooks."

What happens to language when machines start talking back?

In recent times, a friend's social media post sparked curiosity about AI-generated voices. Her dissatisfaction stemmed from the lack of natural breathing and variation in these synthetic voices, which she found to be overly formulaic. This modest concern hints at a more substantial inquiry: how language evolves when machines begin conversing back.

Communication is a mutual endeavor; listeners play a vital role, assessing not just the content but also the delivery, context, and expectations. Even imperfections in speech, often labeled as disfluencies, convey meaning within human interactions. However, synthetic voices that eliminate disfluencies might appear grammatically correct but lack the richness that listeners appreciate.

Language is not merely a sequence of words and pronunciation rules; it is intertwined with memory, personal experiences, and cultural context. This perspective, articulated by author Arundhati Roy, emphasizes the fluid and dynamic nature of human communication. The nuances of language extend beyond mere word exchanges; they involve the listener's memories, expectations, and emotions.

Two individuals can interpret the same sentence differently, and humor is deeply influenced by these factors. Social characteristics of a voice also play a crucial role in communication, a point underscored by linguist Nicole Holliday's research on audience judgments of AI voices. While the focus of generative AI development has largely centered on the system's ability to produce language, the listener's perspective has been overlooked.

As AI systems become integral to our daily interactions, the challenge lies in designing them to be sensitive to human needs and responses. Generative AI systems, though capable of producing language, have been underdeveloped in terms of considering the listeners' experiences. As these systems interact more frequently with humans, they may begin to emulate human-like rhythms and prosody.

However, there is a possibility that these synthetic voices may develop unique linguistic patterns, akin to a watermark, distinguishing them from human speech. The line between flaw and stylistic feature may blur, depending on how society perceives these traits. Innovative language development has historically been driven by human interaction, with communities coining new expressions, borrowing from each other, and evolving the language.

Now, with AI agents generating language at a pace and scale beyond human capabilities, the landscape of language creation is shifting. It is essential to consider how synthetic systems interact with each other and contribute to the linguistic evolution, potentially introducing novel expressions and metaphors. The design of AI systems must prioritize the listener's role, ensuring that AI-generated language respects and incorporates diverse voices, especially those from marginalized groups.

The ultimate question is not merely whether AI sounds human but what kind of human language it reflects, and how this impacts the representation of language for underrepresented communities. The listener remains central to this evolving communication paradigm. As humans grapple with the language produced by machines, the interplay between human and machine communication will continue to shape the future of language.

The outcomes remain uncertain, but the importance of listening and adapting to these changes is clear.

Written by urgent.news from Phys.org's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at phys.org →

More in AI

More from Thursday 8 October →