Google's SL2T brings sign language AI to Pixel phones
In a significant step for accessibility technology, Google DeepMind has introduced SL2T, a multilingual sign language-to-text model, and integrated it into the Pixel 11 smartphone system. This marks the first time sign language AI has been embedded in mainstream consumer electronics, opening up new possibilities for deaf and hard-of-hearing users.
The model is initially integrated with Gboard keyboard and Live Transcribe, enabling users to compose messages, search the web, and interact with Gemini by performing sign language gestures in front of the phone's camera. This replaces the traditional keyboard input, offering a more natural and intuitive way to communicate.
SL2T was trained on an extensive dataset comprising over 100,000 hours of sign language video, covering more than 50 sign languages. About a quarter of this data comes from American Sign Language (ASL) datasets. The cross-lingual joint training approach allows the model to learn common action patterns across different sign languages, and it has set a new record for sign language transcription models on the FLEURS-ASL evaluation benchmark.
Unlike traditional solutions that rely on fixed vocabulary tags, SL2T directly interprets human movements to generate text. It also processes non-hand semantic information such as facial expressions, body position in space, and lip movements, providing a more comprehensive understanding of the signer's intent.
Privacy is a key consideration. The phone-side MediaPipe Holistic tracks 130 key points on the hands, face, and torso in real time, transmitting only movement coordinates externally. The original video is immediately deleted and never uploaded to the cloud, ensuring user privacy.
Currently, the feature supports only American Sign Language to English translation. Google plans to expand to more sign languages and Android devices in the future, making this technology accessible to a wider audience.
DeepMind previously open-sourced a sign language translation model, SignGemma, in May 2025, providing an algorithmic foundation for SL2T. Additionally, its WaveNet, Euphonia, and Live Transcribe accessibility technologies have accumulated experience in multimodal audio and visual interaction.
As Google, iFLYTEK, Huawei, and Apple continue to invest in accessibility, AI-driven interaction is accelerating from research validation to consumer-level products. This development not only showcases technological advancement but also underscores the importance of inclusive design in the digital age.
Key Points
- First of its kind: SL2T is the first sign language AI integrated into consumer smartphones.
- Multilingual training: Trained on 100,000+ hours of data from 50+ sign languages.
- Advanced understanding: Processes facial expressions, body position, and lip movements beyond hand gestures.
- Privacy-focused: On-device processing with immediate deletion of original video.
- Future expansion: Initially ASL to English, with plans for more languages and devices.