Google's AI Language Models Go Beyond Translation to Capture Tone and Slang

Google's AI language efforts now cover more than 300 languages, aiming to understand how people actually speak, including tone and slang. The company is focusing on cultural nuance and working with local communities to ensure tools work even without reliable internet.
Google's Universal Speech Model was trained on 12 million hours of audio and employs cross-lingual transfer learning, letting insights from data-rich languages improve speech recognition in languages with far fewer training resources. The company's 1,000 Languages Initiative targets the world's most-spoken languages beyond those where AI currently performs best, with community partnerships addressing connectivity gaps.
Gemini 3.5 Live Translate provides real-time spoken translation across 70 languages and more than 2,000 language pairs, capturing code-switching and emotional cues. The Transcribe model converts raw audio into formatted text even amid background noise or specialized jargon, while the Rambler feature on Android Gboard removes filler words, corrects grammar, and enables voice-based editing across languages.
This expansion could meaningfully reduce the digital divide for billions of speakers of underrepresented languages, granting broader access to information, education, and services. Communities with limited connectivity may gain practical tools for daily communication, though the systems' ability to preserve genuine cultural nuance remains unproven. The technology could also shape how smaller languages evolve in digital spaces, potentially influencing their long-term vitality and usage patterns.