CrossModalQA:面向多模态检索增强生成的跨模态多跳基准
原标题:CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation
AI 摘要
本文提出了 CrossModalQA,一个用于评估多模态检索增强生成(RAG)的开域基准,包含 1,863 个问答对,基于 4,987 篇维基百科文章和 4,431 张维基共享图片构建,覆盖五种跨模态推理路径,平均推理深度为 3.50 跳。实验表明现有多模态 RAG 系统难以恢复完整证据链,且不完整检索可能引入干扰,导致性能低于闭卷模型;完整跨模态检索比生成器规模对准确率贡献更大,多图像检索与推理是主要瓶颈。
正文节选
CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation Abstract Despite the strong capabilities of multimodal large language models (MLLMs), their parametric knowledge remains incomplete and difficult to update, motivating multimodal retrieval-augmented generation (RAG) to ground responses in external text and images. However, existing benchmarks face two major limitations: (i) they typically emphasize single-hop retrieval or reasoning over a small set