返回全部动态

CrossModalQA:面向多模态检索增强生成的跨模态多跳基准

原标题:CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation

arXiv cs.CV一手来源研究质量 88

AI 摘要

本文提出了 CrossModalQA,一个用于评估多模态检索增强生成(RAG)的开域基准,包含 1,863 个问答对,基于 4,987 篇维基百科文章和 4,431 张维基共享图片构建,覆盖五种跨模态推理路径,平均推理深度为 3.50 跳。实验表明现有多模态 RAG 系统难以恢复完整证据链,且不完整检索可能引入干扰,导致性能低于闭卷模型;完整跨模态检索比生成器规模对准确率贡献更大,多图像检索与推理是主要瓶颈。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation Abstract Despite the strong capabilities of multimodal large language models (MLLMs), their parametric knowledge remains incomplete and difficult to update, motivating multimodal retrieval-augmented generation (RAG) to ground responses in external text and images. However, existing benchmarks face two major limitations: (i) they typically emphasize single-hop retrieval or reasoning over a small set


发布时间:2026-09-09 12:00
抓取时间:2026-09-09 12:28
来源机构:arXiv
阅读原文arxiv.org