Google has unveiled two new versions of its Gemini speech AI system, which it claims will revolutionize the way people interact with digital content. The company is introducing dedicated speech generation systems engineered specifically for direct performance scripting and high-volume audio production.
The two Gemini 3.8 Flash TTS (text-to-speech) voice models are designed to tackle tasks that require creative direction as well as cost management, according to Google. This means studios and content creators can now use the AI system to generate multiple voices with minimal additional expense, reducing labor costs and allowing for more efficient production.
The Gemini 3.8 Flash TTS is specifically aimed at interactive entertainment, game development, and long-form narrations where rapid vocal updates are critical. By breaking down vocal synthesis tasks between creative direction and infrastructure management, the AI system can provide a seamless user experience with minimal manual intervention. This marks an important step forward for Google's Gemini speech AI, which has already demonstrated its potential in a range of applications including voice assistants and chatbots.