返回全部动态

动态潜在推理:先感知后推理的视频理解与问答新方法

原标题:Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering

arXiv cs.CV一手来源研究质量 84

AI 摘要

arXiv 上发布了一项关于视频问答的新研究,提出了动态潜在推理(DyLaR)方法。该方法先通过感知潜在状态定位问题相关的视觉证据,再自适应决定是否在潜在空间中进行推理,从而避免冗长的文本思维链。在九个视频基准和四个多模态语言模型上的实验显示,DyLaR 在提升准确率的同时大幅减少了生成的 token 数量,例如在 Qwen3-VL-4B 上将平均准确率从 54.0 提升至 58.2,响应长度从 1220.7 降至 18.5 token。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Computer Vision and Pattern Recognition Title:Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering View PDF HTML (experimental) Abstract:Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as soon as the re


发布时间:2026-08-06 12:00
抓取时间:2026-08-06 21:38
来源机构:arXiv
阅读原文arxiv.org