返回全部动态

Meta 发布实时音频模型 Muse Voice Transcribe,支持多说话人识别

原标题:Meta's new real-time audio model is the foundation for AI assistants that never stop listening

THE DECODER模型发布质量 72

AI 摘要

Meta 的超级智能实验室发布了首个实时音频感知模型 Muse Voice Transcribe,能够实时转写语音、区分说话人并检测句子边界。该模型将音频分成 80 毫秒的块,并根据难度动态调整延迟,以平衡速度和准确性。其定价为每小时 0.18 美元,低于 OpenAI 和 ElevenLabs 等竞争对手,支持超过 70 种语言,现已通过 Meta AI 和 Meta Model API 提供。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Meta's new real-time audio model is the foundation for AI assistants that never stop listening Key Points - Meta has released Muse Voice Transcribe, a real-time model that transcribes speech, detects sentence boundaries, and tells up to 20 speakers apart without separate systems. - The model breaks audio into 80-millisecond chunks and adjusts the delay for each word based on difficulty, balancing speed and accuracy on the fly. - At $0.18 per hour, Meta undercuts competitors like OpenAI and Eleve


发布时间:2026-09-06 17:45
抓取时间:2026-09-06 18:32
来源机构:THE DECODER
阅读原文the-decoder.com