零令牌置信度:读取模型内部状态,0.06秒零令牌完成验证
原标题:Your model already knows it's wrong. Asking costs 0.06 seconds and zero tokens.
AI 摘要
VIDRAFT 发布零令牌置信度(ZTC)方法,通过读取模型内部隐藏状态而非生成令牌来评估答案可信度,单次前向传播即可输出校准概率,已在 Hugging Face 上线。在 2,018 项共享基准上,ZTC·Darwin-397B 以 0.7394 AUC 登顶,而模型自报置信度仅 0.5000(等同随机猜测)。在 539 项留出集上,读取内部状态比直接询问模型提升 0.116 至 0.225 AUC,且门控耗时仅 0.0615 秒,比其守护的生成工作便宜 26 倍。
正文节选
Read its internal state instead and you get 0.8801. Same model. Same questions. The verdict was already there; nobody was reading it. Zero-Token Confidence (ZTC) reads it. One forward pass, zero generated tokens, a calibrated probability out the other side. It ships today on Hugging Face. A shared benchmark, 2,018 items, one harness, every verifier scored on the same rows. | Verifier | AUC | Generated tokens | |---|---|---| | 🥇 ZTC · Darwin-397B | 0.7394 | 0 | | JEV | 0.7335 | — | | 🥉 ZTC-Jud