字节跳动发布 Seed Audio 1.0 场景级音频生成模型
原标题:From Speech to Audio Creation | Introducing the Seed Audio 1.0 Audio Creation Model
AI 摘要
字节跳动旗下 Seed Research 团队发布了 Seed Audio 1.0 音频生成模型,该模型能在统一框架内生成语音、音效、环境音等场景级音频元素,支持 20 多种语言、最长两分钟的单次生成,并具备 100 毫秒级别的对话时序控制能力。该模型旨在帮助创作者从拼接音频片段转向导演式的声音场景创作,提升叙事音频、视频配音、游戏本地化等场景的制作效率。
正文节选
From Speech to Audio Creation | Introducing the Seed Audio 1.0 Audio Creation Model From Speech to Audio Creation | Introducing the Seed Audio 1.0 Audio Creation Model Date 2026-07-20 Category Models Full-scene audio generation for creators Creators rarely imagine sound as a set of isolated files. They hear the room before a line is spoken. They hear the pressure in a pause, the footstep just outside the frame, the alarm buried under the dialogue, the background that makes a place feel real, and