MUSES:前瞻性知识根源检索基准
原标题:MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval
AI 摘要
MUSES 是一个面向前瞻性知识根源检索的百万级基准,基于 2.33M 篇论文语料,包含约 14 万测试实例,按熟悉度分为三层。配套的 CiteRoots 框架提供修辞层和作者认可层,其中作者认可层包含 753 篇焦点论文的 1518 对生成灵感对。实验表明,基于 SPECTER2 的多质心检索器表现最佳,但性能随难度下降,且修辞角色与作者认可存在差异。该基准旨在推动前瞻性检索研究。
正文节选
MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval Abstract Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce MUSES, a million-instance benchmark for prospective intellectual-roots retrieval over a fixed 2.33M-paper corpus, with roughly 140K test instances per familiarity tier. To our kn