返回全部动态

从六个日志重建 Among AIs:LLM 推理欺骗但无法执行

原标题:Rebuilding Among AIs from Six Log Files

Hugging Face Blog一手来源研究质量 87

AI 摘要

Antim Labs 发布了 Among AIs 社交推理基准测试,让 LLM 玩《Among Us》游戏。作者从六个日志文件重建了游戏引擎和地图,并用六个小型开放模型运行了 90 局游戏。结果显示,这些模型能推理欺骗但无法执行,船员胜率高达 80%。作者还开源了重建工具和模拟器。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

TL;DR — Antim Labs published Among AIs, a social-deduction benchmark where LLMs play a game of Among Us. They released game logs but not the engine or the map. I reconstructed both from six log files, rebuilt the whole game from scratch, and ran 90 games across six small open models (27B–35B class). The headline result is not the leaderboard. It is this: these models can reason about deception but cannot execute it. Crewmates won 80% of games. One model literally typed its own deception plan int


发布时间:—
抓取时间:2026-08-08 04:55
来源机构:Hugging Face
阅读原文huggingface.co