返回全部动态

ToolHazard:扩展对抗环境以评估和提升LLM代理安全性

原标题:ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

ToolHazard 是一个可扩展的对抗性环境合成框架,用于测试 LLM 代理对间接提示注入的脆弱性。该框架通过环境模拟器、攻击代理和用户模拟器生成可执行的状态化环境,并构建了 ToolHazard-Bench 基准。实验表明代理存在显著脆弱性,且注入时机和位置影响攻击效果。ToolHazard 生成的对抗数据能提升代理在 ToolHazard-Bench 和 AgentDojo 上的安全性,同时保持良性任务效用。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Abstract ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment. Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually


发布时间:—
抓取时间:2026-08-13 09:49
来源机构:Hugging Face
阅读原文huggingface.co