Urgent.News

What's breaking now, across thousands of outlets.

Tech

How do Siri and Alexa sound so human? A computer engineer explains

Curious Kids is a series for children of all ages. If you have a question you’d like an expert to answer, send it to CuriousKidsUS@theconversation.com . How is a text-to-speech voice made? – Sarah G٫ age 11٫ Seguin٫ Texas When you talk to computerized assistants like Siri or Alexa , they reply in voices that sound very human. But how do computers, smartphones and apps actually talk like a person?…

How do Siri and Alexa sound so human? A computer engineer explains

Text-to-speech voices that sound human are created through a process called text-to-speech technology. When you converse with digital assistants like Siri or Alexa, their replies come across as human-like due to this technology. The process begins with the lungs pushing air up the windpipe and through the vocal cords, causing them to vibrate and form sound. Your brain then directs your mouth, tongue, and lips to shape this sound into words.

As a computer engineer specializing in creating realistic experiences for people, I will explain how computers emulate this speech-making process. A computer simulates the creation of spoken words by sending electrical signals to a small speaker, which vibrates rapidly. These rapid vibrations push against the surrounding air, generating sound waves. The computer's software controls these electrical signals to shape the sound waves and produce speech.

To create specific words, the computer breaks them down into tiny sound pieces called phonemes. Phonemes are the smallest units of speech, such as 'sh', 'short i', or 'p' used in the word "ship." The software generates these phonemes and arranges them in the right order to form complete words like "Hello, how are you?"

The pursuit of artificial speech can be traced back to the 1700s, when inventors attempted to make machines function similarly to the human lungs and throat. Early attempts, such as bellows and pipes, produced squeaky, strange, and unsettling sounds. The first electronic speech machines, known as synthesizers, emerged in the 1930s, with a notable example being the Voder, showcased at the 1939 World's Fair in New York City.

Voder resembled an organ and required electronic buttons, keys, and foot pedals to emit basic phrases like "Good morning!"

In the 1960s, computers started constructing speech by combining phonemes, resulting in stiff and robotic-sounding voices, such as "He-llo-hu-man-I-am-a-com-pu-ter." This occurred because older software programs had to combine small sounds mapped from recorded voices, creating a puzzle-like structure that translated those maps back into speech. While this method worked, it produced unnatural-sounding results.

Today's computers employ machine learning, a type of artificial intelligence, to generate speech that closely resembles human voices. Researchers train AI programs using hours of real people's recordings. The machine learning software meticulously analyzes patterns in speech, including breathing, laughter, and variations in voice pitch during excitement. By understanding these patterns, the computer can shape phonemes into coherent words and sentences, capturing the nuances of human speech.

Advanced technology even enables sophisticated AI computers to listen to a few seconds of your voice recording, learn your unique speech patterns, and mimic your voice. Consequently, these AI systems can generate sentences you have never spoken, all while maintaining a voice that sounds remarkably like yours.

Text-to-speech software has a wide range of practical applications, benefiting millions of people daily. In cars, it provides navigational guidance, allowing drivers to maintain focus on the road. It can read news articles and websites for individuals who are blind and even provide a voice for those unable to speak. However, like other powerful technologies, text-to-speech software can also be misused.

Advanced software can produce highly realistic synthetic voices, sometimes referred to as audio deepfakes, which closely resemble real people's voices by analyzing short samples of their speech. Scammers can exploit this technology to impersonate family members, coworkers, or celebrities, potentially deceiving people with false information.

Scientists and engineers are actively developing tools to identify fake voices and combat scams. So, the next time you encounter a phone, computer, or video game speaking with a human-like voice, you now understand the process behind it, without the need for actual lungs, vocal cords, lips, or a tongue!

Written by urgent.news from Fast Company's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at fastcompany.com →

More in Tech

Faked photos could be shown up by new iPhone feature

Apple aims to make it possible to prove that a photo was taken on your iPhone and, presumably, not subsequently altered. The latest developer beta of iOS 27 has an unusual number of new features for…

  • New iOS 27 feature called Apple Reference Image
  • Authenticates photos taken by iPhone camera
  • Relies on specific conditions to prove origin

More from Wednesday 12 August →