返回全部动态

41年Jeopardy!知识时间胶囊:单个免费本地模型的表现

原标题:Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

arXiv cs.AI一手来源研究质量 84

AI 摘要

该研究首次在完整的Jeopardy!题库(1984-2025年,529,939条线索)上评估了单个9GB开源模型(Qwen2.5-14B,4位量化),严格强制回答下准确率达67.0%,事实类类别超过85%。研究对比了IBM Watson,指出训练数据暴露是两者共有的,但本地模型在训练截止后的线索上仍保持65%准确率,而Watson为0%。作者认为知识压缩能力已从服务器机房转移到便携文件,重新定义了Watson挑战。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model Abstract In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy. We show that the same kind of artifact, a snapshot of what a cultu


发布时间:2026-09-01 12:00
抓取时间:2026-08-31 12:10
来源机构:arXiv
阅读原文arxiv.org