TestHallVQA:面向科学考试的LVLM文档级冗余上下文推理基准
原标题:TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams
AI 摘要
研究者提出 TestHallVQA,一个面向大型视觉语言模型(LVLM)的多图像文档级 VQA 基准,数据来自科学考试试卷,包含超过 1 万条问答对和 7000 张试卷图像。该基准可控制地注入多级冗余上下文,并配套提出 F1-R2 指标,同时衡量模型的推理能力与对文档级冗余的证据检索鲁棒性。实验显示主流 LVLM 在多个维度存在潜在缺陷,为后续研究提供方向。
正文节选
TestHallVQA: Exploring LVLMs’ Document-Level Reasoning under Redundant Contexts from Scientific Exams Abstract Large Vision–Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasoning depth, while others require complex visual reasoning but remain restricted to single-page, noise-free settings. Moreo