SIGNPOST-Bench:多模态大语言模型文本-视觉冲突解决基准
原标题:SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
AI 摘要
SIGNPOST-Bench 是一个用于评估多模态大语言模型(MLLMs)在文本与视觉信息冲突时仲裁能力的基准。该基准通过将图像转换为原始、空白、相似、随机和对抗五种变体,构建了 5,111 个反事实组和 25,555 个图像变体,并评估了来自七个提供商的 20 个 MLLM。结果显示,对抗性变体使中位定位误差从 282 公里增加到 1,347 公里,提升了 4.8 倍,且所有模型均表现出对冲突文本的脆弱性。该研究为评估 MLLMs 解决多模态证据冲突提供了受控框架。
正文节选
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Abstract Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a count