Gemini 3.1 Flash TTS 发布:更自然、可控的 AI 语音生成
原标题:Gemini 3.1 Flash TTS: the next generation of expressive AI speech
AI 摘要
Google DeepMind 发布了 Gemini 3.1 Flash TTS,这是一款新一代文本转语音模型,提升了可控性、表现力和质量。该模型已通过 Gemini API、Google AI Studio、Vertex AI 和 Google Vids 向开发者、企业和 Workspace 用户推出,支持 70 多种语言,并引入了音频标签功能,允许通过自然语言指令精细控制语音风格、语速和语调。模型在 Artificial Analysis TTS 排行榜上获得 1211 的 Elo 评分,所有生成音频均带有 SynthID 水印,以防止滥用。
正文节选
Gemini 3.1 Flash TTS: the next generation of expressive AI speech Today, we’re introducing Gemini 3.1 Flash TTS, the latest text-to-speech model that delivers improved controllability, expressivity and quality — empowering developers, enterprises and everyday users to build the next generation of AI-speech applications. Starting today, 3.1 Flash TTS is rolling out: - For developers in preview via the Gemini API and Google AI Studio - For enterprises in preview on Vertex AI - For Workspace users