动态潜在推理:先感知后推理的视频理解与问答新方法
原标题:Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
AI 摘要
arXiv 上发布了一项关于视频问答的新研究,提出了动态潜在推理(DyLaR)方法。该方法先通过感知潜在状态定位问题相关的视觉证据,再自适应决定是否在潜在空间中进行推理,从而避免冗长的文本思维链。在九个视频基准和四个多模态语言模型上的实验显示,DyLaR 在提升准确率的同时大幅减少了生成的 token 数量,例如在 Qwen3-VL-4B 上将平均准确率从 54.0 提升至 58.2,响应长度从 1220.7 降至 18.5 token。
正文节选
Computer Science > Computer Vision and Pattern Recognition Title:Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering View PDF HTML (experimental) Abstract:Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as soon as the re