基于扰动的 VQA 图像区域标注方法 CSGR
原标题:Don't Just Look, Intervene: Perturbation Based Region Labeling for VQA Images
AI 摘要
该论文提出了一种名为 CSGR(Counterfactual Search for Grounding Regions)的自动标注流程,通过分割模型提出候选区域、用修复模型扰动这些区域,并测量其对模型答案分布的影响,从而识别出对答案起关键作用的「模型因果视觉证据」区域。CSGR 在多个评判模型间聚合证据以减少偏差,生成细粒度像素级标注。实验将该标注接入注意力引导、Visual CoT 微调和潜在视觉推理三种训练流程,结果显示其在内域和外域评估中均比仅用交叉熵微调取得更一致的提升。
正文节选
Don’t Just Look, Intervene: Perturbation Based Region Labeling for VQA Images Abstract Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, defined as the set of image regions whose counterfactual intervention changes a model’s answer d