Google debuts SL2T, an AI model that’s designed to understand sign language
Google DeepMind said today it wants to bring the artificial intelligence revolution to the estimated 70 million people across the world who are either deaf or hard of hearing with the launch of sign-language-to-text or SL2T. In a blog post, Google’s AI researchers said SL2T is a multilingual translation model that’s making its debut on […] The post Google debuts SL2T, an AI model that’s designed…
Google DeepMind has unveiled an innovative AI model called SL2T, which stands for sign-language-to-text. Aimed at assisting the estimated 70 million individuals worldwide who are either deaf or hard of hearing, this new technology translates sign language into text, making its debut on the Pixel 11 smartphone. SL2T is a multilingual translation model capable of converting American Sign Language (ASL) into English text, and it marks the first sign language translation feature integrated into a consumer product.
Prior to this, AI had advanced significantly in processing spoken human speech, enabling users to dictate messages in multiple languages on various devices. However, sign language users, who rely on one of the 200+ distinct sign languages, have been largely overlooked by the AI industry until now.
The SL2T model empowers deaf and hard-of-hearing users to interact with their smartphones using their native sign language. This capability enables users to perform various tasks, such as web searches, drafting emails, and editing documents, all through signing. Users can also utilize SL2T to prompt Google's Gemini AI to answer questions and perform actions on their behalf. The model is available in Google's Live Transcribe app, allowing users to communicate via sign language during face-to-face calls.
SL2T was trained on over 100,000 hours of sign language data spanning 50 languages, with approximately a quarter of the dataset dedicated to ASL communications. Although the initial release focuses solely on ASL, Google decided to train the model across multiple sign languages to learn the shared structural patterns, enabling it to outperform earlier sign language models.
To address privacy concerns, SL2T utilizes an on-device computer vision model called MediaPipe Holistic, which tracks and sends the coordinates of the signer's facial, hand, arm, and torso movements to a cloud-based server, bypassing the need to upload actual video.
By directly translating sign language sequences into text without intermediate text annotations, SL2T better captures non-manual expressions and spatial grammar structures specific to ASL. Additional features include latency optimizations, mechanisms to prevent hallucinations for non-signing movements, and support for left-handed signers and one-handed signing. Google's performance claims are supported by solid data, with SL2T achieving a 70 BLEURT score on the FLEURS-ASL benchmark.
This groundbreaking development is crucial for the global deaf community, as sign language understanding has been a long-forgotten area in AI research. Google has established an AI Sign Language Advisory Committee, composed of deaf organizations and sign language experts, to guide the responsible deployment of this technology. The committee will also co-author a report detailing SL2T's capabilities and limitations.
Looking ahead, Google plans to expand SL2T to cover additional sign languages and develop models for sign language generation.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.