返回全部动态

Gemini 3.1 Flash TTS 发布:更自然、可控的 AI 语音生成

原标题:Gemini 3.1 Flash TTS: the next generation of expressive AI speech

Google DeepMind News一手来源模型发布质量 80

AI 摘要

Google DeepMind 发布了 Gemini 3.1 Flash TTS,这是一款新一代文本转语音模型,提升了可控性、表现力和质量。该模型已通过 Gemini API、Google AI Studio、Vertex AI 和 Google Vids 向开发者、企业和 Workspace 用户推出,支持 70 多种语言,并引入了音频标签功能,允许通过自然语言指令精细控制语音风格、语速和语调。模型在 Artificial Analysis TTS 排行榜上获得 1211 的 Elo 评分,所有生成音频均带有 SynthID 水印,以防止滥用。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Gemini 3.1 Flash TTS: the next generation of expressive AI speech Today, we’re introducing Gemini 3.1 Flash TTS, the latest text-to-speech model that delivers improved controllability, expressivity and quality — empowering developers, enterprises and everyday users to build the next generation of AI-speech applications. Starting today, 3.1 Flash TTS is rolling out: - For developers in preview via the Gemini API and Google AI Studio - For enterprises in preview on Vertex AI - For Workspace users


发布时间:2026-04-16 00:03
抓取时间:2026-09-07 04:18
来源机构:Google DeepMind
阅读原文deepmind.google