返回全部动态

Google 发布 Gemini 3.5 Transcribe 高精度语音转文字模型

原标题:Intelligent transcription with Gemini 3.5 Transcribe

Google DeepMind News一手来源模型发布质量 82

AI 摘要

Google DeepMind 发布了 Gemini 3.5 Transcribe,一款高精度语音转文字模型,支持实时流式与预录音频处理,具备智能转录、函数调用、自定义词汇、85+ 语言及多说话人识别等功能。该模型已集成到 Gemini API、Google AI Studio 及 Gemini Enterprise Agent Platform,并应用于 Gboard、Gemini 应用等产品。相比前代 Chirp 3,其词错误率显著降低,转录速度提升 70%,开发者可通过 Live API 和 Interactions API 使用。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Intelligent transcription with Gemini 3.5 Transcribe Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting f


发布时间:2026-08-27 01:01
抓取时间:2026-09-07 08:45
来源机构:Google DeepMind
阅读原文deepmind.google