返回全部动态

LitReview Arena:基于竞技场式同行评审平台的文献综述智能体评估

原标题:LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

arXiv cs.AI一手来源研究质量 85

AI 摘要

LitReview Arena 是一个用于评估文献综述智能体的竞技场式平台,通过领域专家对匿名草稿进行多维度的偏好投票。研究收集了约3000条专家判断,发现最强模型在整体效用上仅以23.0%的胜率击败人类草稿,而Sonar Deep Research等智能体LLM比基础模型性能提升超过60%。此外,现有LLM-as-a-judge方法与人类专家存在显著不一致,研究者因此提出了专家校准的评估器LitJudge,其对齐度接近专家间一致性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform Abstract Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-


发布时间:2026-08-25 12:00
抓取时间:2026-08-25 12:02
来源机构:arXiv
阅读原文arxiv.org