返回全部动态

Gemini 3.5 Transcribe 模型全面上市,支持流式与非流式语音转文字

原标题:August 26, 2026

Gemini API Release Notes一手来源模型发布质量 79

AI 摘要

Gemini API 于 2026 年 8 月 26 日发布 Gemini 3.5 Transcribe 和 Gemini 3.5 Transcribe Live 两款语音转文字模型,均基于 Gemini 的音频理解能力。前者为非流式模型,支持 85+ 语言检测、说话人分离、词级时间戳和自定义词汇偏置;后者为低延迟双向流式模型,支持中间和最终转录事件、智能转录模式及多种 VAD 策略。这些模型已全面可用,开发者可通过相关指南和模型页面开始使用。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Gemini 3.5 Transcribe generally available (GA): Released two dedicated speech-to-text models based on Gemini's audio understanding: Gemini 3.5 Transcribe ( gemini-3.5-transcribe ): High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps, and custom vocabulary biasing (up to 1,000 terms). Gemini 3.5 Transcribe Live ( gemini-3.5-transcribe-live ): Low-latency, bidirectional streaming speech-to-


发布时间:2026-08-26 08:00
抓取时间:2026-08-31 00:41
来源机构:Google
阅读原文ai.google.dev