Omni-Streaming Thinking:延迟跨模态验证的流式全模态推理
原标题:Paper page - Omni-Streaming Thinking
AI 摘要
该论文提出 Omni-Streaming Thinking(OST)方法,用于改进流式全模态推理。针对模型在音频尚未完整时过早将视觉解读固化为事实、导致后续即使音频矛盾仍持续传播的「跨模态过早承诺」问题,OST 生成包含已观察证据、未来证据预测和基于证据的声明等结构化输出,声明初始标记为待定并绑定未来验证区间,音视频证据分开存储,在验证区间结束时进行跨模态核验,发现矛盾则进行反驳并更新状态,再由答案门控决定是否作答。基于冻结的 Qwen3-Omni-30B-A3B-Instruct 主干加轻量适配,OST 在五个流式与音视频基准上平均相对提升超过 10%,并引入 OST-DiagBench 诊断基准,d-prime 达 2.95(开放基线最高 1.38),同时减少视觉诱发的听觉幻觉。
正文节选
Omni-Streaming Thinking Abstract Omni-Streaming Thinking improves streaming omni-modal reasoning by deferring claims until cross-modal verification, reducing premature commitment and auditory hallucinations. Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep r