返回全部动态

基于扰动的 VQA 图像区域标注方法 CSGR

原标题:Don't Just Look, Intervene: Perturbation Based Region Labeling for VQA Images

arXiv cs.CV一手来源研究质量 78

AI 摘要

该论文提出了一种名为 CSGR(Counterfactual Search for Grounding Regions)的自动标注流程,通过分割模型提出候选区域、用修复模型扰动这些区域,并测量其对模型答案分布的影响,从而识别出对答案起关键作用的「模型因果视觉证据」区域。CSGR 在多个评判模型间聚合证据以减少偏差,生成细粒度像素级标注。实验将该标注接入注意力引导、Visual CoT 微调和潜在视觉推理三种训练流程,结果显示其在内域和外域评估中均比仅用交叉熵微调取得更一致的提升。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Don’t Just Look, Intervene: Perturbation Based Region Labeling for VQA Images Abstract Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, defined as the set of image regions whose counterfactual intervention changes a model’s answer d


发布时间:2026-09-15 12:00
抓取时间:2026-09-15 12:22
来源机构:arXiv
阅读原文arxiv.org