返回全部动态

Ex-Omni-2D:具有原生视觉呈现的全模态对话模型

原标题:Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Hugging Face Daily Papers一手来源研究质量 83

AI 摘要

Ex-Omni-2D 是一个全模态对话框架,能够生成协调的文本、个性化语音和参考条件视频响应。该模型通过视觉思维计划和蒸馏的流式视频生成器,实现了视觉呈现的对话能力。其共享声学-时间接口支持从异构数据中学习,避免了大规模查询-文本-语音-视频监督的需求。在四步推理下,完整的四 GPU 流水线实现了 1.293 的端到端 RTF,提供了实用的质量-效率平衡点。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Abstract Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator. Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response c


发布时间:—
抓取时间:2026-08-12 12:13
来源机构:Hugging Face
阅读原文huggingface.co