返回全部动态

自验证智能体工具:分离长时程智能体中的承诺漂移与绑定漂移

原标题:The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

arXiv cs.AI一手来源研究质量 82

AI 摘要

该研究提出一种自验证的智能体工具,通过确定性执行器控制信念、语言模型仅提交类型化提案,并利用预注册预测与代码观测匹配来验证长时程智能体。实验表明,移除承诺机制会导致目标放弃率从0.00升至1.00,而绑定错误保持为0.00,验证了漂移分解的有效性。尽管任务效能为零(52次门控运行中无一次完成ARC-AGI-3),但该方法为智能体开发提供了结构性验证方法论。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Artificial Intelligence Title:The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents View PDF HTML (experimental) Abstract:How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive owns all belief; a


发布时间:2026-08-06 12:00
抓取时间:2026-08-06 21:15
来源机构:arXiv
阅读原文arxiv.org