幻觉作为特征而非缺陷:多智能体架构将推测性语言模型输出转化为可测试科学假设的评估
原标题:Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
AI 摘要
本文提出一种基于Rust的多智能体架构,通过将高熵生成代理与基于网络搜索的评估代理分离,利用“认识论摩擦”将语言模型的幻觉输出转化为可测试的科学假设。实验表明,直接提示效果最差,但完整系统并不总是优于简单自我反思,其优势主要体现在假设需满足强物理、经验或制度约束时。研究强调,受控的推测性生成只有在架构、经验基础和显式评估的约束下才有价值。
正文节选
Hallucination as a Feature, not a Defect Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses Abstract Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also restrict speculative Research and Development (R&D) by encouraging what this work operationally treats