返回全部动态

字节发布 SeedRealtime 音视频全双工大模型,实现多模态自然交互

原标题:SeedRealtime Audio-Visual Full-Duplex LLM Released: Toward Omni-Modal Natural Interaction

ByteDance Seed Research一手来源模型发布质量 86

AI 摘要

字节跳动旗下 Seed Research 正式发布 SeedRealtime,一款原生音视频全双工大语言模型。该模型采用统一架构原生融合音频、视频和文本,实现实时多模态交互,具备联合音视频理解、主动交互和自然对话节奏三大突破。端到端评估显示,相比级联模型,其音视频对话节奏问题减半,并已在行业率先实现大规模部署。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Today, we are officially launching SeedRealtime, a native audio-visual full-duplex LLM. As a key step toward omni-modal interaction, SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen, and speak" experience. It achieves three core breakthroughs: Joint audio-visual understanding: Native support for the deep fusion of audio, visual, and temporal information. Th


发布时间:—
抓取时间:2026-08-17 03:43
来源机构:ByteDance Seed
阅读原文seed.bytedance.com