Ex-Omni-2D:具有原生视觉呈现的全模态对话模型
原标题:Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
AI 摘要
Ex-Omni-2D 是一个全模态对话框架,能够生成协调的文本、个性化语音和参考条件视频响应。该模型通过视觉思维计划和蒸馏的流式视频生成器,实现了视觉呈现的对话能力。其共享声学-时间接口支持从异构数据中学习,避免了大规模查询-文本-语音-视频监督的需求。在四步推理下,完整的四 GPU 流水线实现了 1.293 的端到端 RTF,提供了实用的质量-效率平衡点。
正文节选
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Abstract Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator. Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response c