How do computers talk with voices that sound like people?
Computers break examples of human speech into small bits of sound, then put them together to sound like people do.
You might have heard phones, computers, or video games speaking with a voice that sounds almost like a human. Curious about how this happens? Sarah G., age 11 from Seguin, Texas, asked this question. Computers use a technology called text-to-speech to replicate human speech. When we talk, air from our lungs travels up the windpipe and through the vocal cords in our throat, causing them to vibrate and create sound.
Our brains then direct our mouth, tongue, and lips to shape that sound into words. A computer simulates this speech creation process by sending electrical signals to a tiny speaker that vibrates rapidly. These vibrations generate sound waves, which the computer controls to mimic speech. To form a speech, a computer breaks words into phonemes – the smallest units of speech, like "sh," "short i," and "p" in "ship."
The computer then groups these phonemes in the correct order to create words. The concept of creating human-like computer speech has been around since the 1700s, with early attempts using mechanical devices that produced squeaky and unnatural sounds. The first electronic speech machine, called Voder, debuted at the 1939 World’s Fair in New York City and resembled an organ.
By the 1960s, computers began speaking by combining phonemes, resulting in stiff and robotic-sounding voices. Today, computers utilize machine learning – a form of artificial intelligence – to create speech that sounds more human-like. Engineers and scientists train AI programs using hours of real people's recorded speech, analyzing patterns in their speech, such as breathing, laughter, and voice changes.
This allows the computer to shape phonemes into words and sentences in a way that closely resembles natural human speech. This advanced technology even enables sophisticated AI computers to learn a person's voice from just a few seconds of recorded speech and mimic it. Text-to-speech software has proven helpful in various aspects of daily life, such as providing directions in cars, reading news articles for the blind, and giving voices to people who cannot speak.
However, it also poses risks, as advanced software can create highly realistic fake voices, or audio deepfakes, which can be used for malicious purposes like impersonating individuals for fraudulent activities. Scientists and engineers are working on identifying fake voices to combat these scams. So, next time you hear a phone, computer, or video game speak with a human voice, you now know how it achieved this without having lungs, vocal cords, lips, or a tongue!
Written by urgent.news from The Conversation's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.