Gemini 2.5 Flash Native Audio 升级,新增实时语音翻译
原标题:Improved Gemini audio models for powerful voice experiences
AI 摘要
Google DeepMind 发布了升级版 Gemini 2.5 Flash Native Audio 模型,用于实时语音代理,提升了函数调用、指令遵循和对话流畅性。该模型已在 Google AI Studio、Vertex AI 上线,并开始集成到 Gemini Live 和 Search Live。同时,Google 推出了实时语音翻译功能,支持 70 多种语言和 2000 种语言对,已在 Google Translate 应用中开启测试。
正文节选
Improved Gemini audio models for powerful voice interactions Earlier this week, we introduced greater control over audio generation with an upgrade to our Gemini 2.5 Pro and Flash Text-to-Speech models. But generating expressive speech is only one side of the conversation. Today, we’re releasing an updated Gemini 2.5 Flash Native Audio for live voice agents. This update improves the model’s ability to handle complex workflows, navigate user instructions, and hold natural conversations. Gemini 2.