返回全部动态

间接 tipping:AI 智能体群体中的社会攻击面

原标题:Indirect tipping: a social attack surface in AI agent populations

arXiv cs.MA一手来源研究质量 88

AI 摘要

该研究指出,评估AI智能体群体安全性时常用的临界质量框架存在低估风险,因为它只关注直接竞争推翻均衡所需的最小对抗比例。作者通过大语言模型智能体群体实验和解析框架,发现经由中间「踏脚石」均衡的间接 tipping 可显著降低所需少数派规模,绕过多数要求,实现直接挑战无法达成的状态转换。结果表明均衡的抗干预性并非内在属性,而是其与替代状态竞争关系的结构性特征,保障多智能体系统安全需同时映射这一社会景观。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Indirect tipping: a social attack surface in AI agent populations Abstract As generative AI agents are deployed at scale, safety will increasingly depend not only on technical safeguards and the design of individual models, but also on the collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same equilibria that enable agents to coordinate also create a social attack surface. The standard framework to assess this


发布时间:2026-09-23 12:00
抓取时间:2026-09-23 12:09
来源机构:arXiv
阅读原文arxiv.org