返回全部动态

OpenART:通过开放式环境演化扩展智能体红队测试

原标题:OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

OpenART 是一个用于可扩展智能体红队测试的开放竞技场,通过环境演化来评估长周期 AI 智能体的安全性。它提供了超过 10,000 个跨 50 个领域的验证状态场景,并提出了 EMHA 攻击策略,该策略在无需参数更新的情况下实现了 85.0% 的总体攻击成功率。研究表明,随着任务复杂度的增加,环境演化能更有效地暴露安全失败,且智能体的运行时实现对其安全性的影响显著。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Abstract OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows. AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a s


发布时间:—
抓取时间:2026-08-13 12:59
来源机构:Hugging Face
阅读原文huggingface.co