知识综合综述框架:基于任务的LLM系统多源证据综合基准测试
原标题:Knowledge Synthesis Review Framework: Task-Level Benchmarking of LLM-Based Systems for Multi-Source Evidence Synthesis
AI 摘要
本文提出知识综合综述(KSR)框架,将证据综合分解为筛选、提取、分析和综合四个任务,并在244篇文档基准上评估GPT-5、Claude Sonnet 4、Gemini 2.5 Pro和NotebookLM。结果显示没有系统在所有任务上领先,Claude在筛选准确率最高,GPT-5召回率最高但特异性较低,解释性分析和跨源综合仍需专家判断。该框架提供透明、可审计、模型无关的LLM辅助研究综合治理方式。
正文节选
Knowledge Synthesis Review Framework: Task-Level Benchmarking of LLM-Based Systems for Multi-Source Evidence Synthesis Abstract Evidence in rapidly evolving fields is fragmented across academic studies, industry reports, policy documents, and media sources that differ in quality, structure, and purpose, making timely synthesis difficult. Large language models (LLMs) may accelerate this work, but their reliability across the distinct cognitive tasks of a review remains uncertain. We introduce the