返回全部动态

AgentGuard:从异常编码智能体轨迹学习执行护栏

原标题:AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories

arXiv cs.SE一手来源研究质量 80

AI 摘要

约克大学Lassonde工程学院的研究者提出AgentGuard,一个从异常编码智能体轨迹中自动学习执行护栏的指令级框架。该方法从642条真实失败轨迹中提取可复用的执行约束,并组织为轻量级技能,仅在触发条件满足时动态激活。使用Claude Code与Claude Haiku 4.5在100个任务上评估,异常执行率从69.0%降至26.7%,任务成功率从21.7%升至35.0%,同时指出过度拒绝是主要局限。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories Abstract AI coding agents increasingly rely on execution harnesses to interact with repositories and external tools. However, task success does not guarantee reliable execution. Agents may still modify unrelated files, rewrite tests, issue unsafe commands, or ignore failed validations, motivating behavioral guardrails for reliable execution. We present AgentGuard, an instruction-level guardrail framework that lear


发布时间:2026-09-16 12:00
抓取时间:2026-09-16 12:52
来源机构:arXiv
阅读原文arxiv.org