空架还是丢钥匙?回忆是参数化事实性的瓶颈
原标题:Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
AI 摘要
Google Research 的研究人员提出知识剖析框架,发现前沿大模型(如 Gemini3 和 GPT-5)的事实编码接近饱和,但回忆能力不足,导致许多事实错误源于无法检索而非未编码。他们构建了包含 2150 个事实的 WikiProfile 基准,并评估了 13 个模型,结果表明扩展模型规模主要改善编码而非回忆,且长尾事实和反向问题的瓶颈在于利用而非获取。
正文节选
August 12, 2026 Nitay Calderon and Gal Yona, Research Scientists, Google Research When LLMs get facts wrong, is it because they never learned them or because they can't recall what they’ve already encoded? Our knowledge profiling framework reveals the latter: frontier LLMs encode nearly all facts, yet struggle to recall many of them. Factuality is essential for making Large Language Models (LLMs) reliable. When a model answers a factual question incorrectly, is it because the fact was never enco