返回全部动态

SIGNPOST-Bench:多模态大语言模型文本-视觉冲突解决基准

原标题:SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

SIGNPOST-Bench 是一个用于评估多模态大语言模型(MLLMs)在文本与视觉信息冲突时仲裁能力的基准。该基准通过将图像转换为原始、空白、相似、随机和对抗五种变体,构建了 5,111 个反事实组和 25,555 个图像变体,并评估了来自七个提供商的 20 个 MLLM。结果显示,对抗性变体使中位定位误差从 282 公里增加到 1,347 公里,提升了 4.8 倍,且所有模型均表现出对冲突文本的脆弱性。该研究为评估 MLLMs 解决多模态证据冲突提供了受控框架。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Abstract Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a count


发布时间:—
抓取时间:2026-08-07 00:21
来源机构:Hugging Face
阅读原文huggingface.co