完成任务还不够:评估累积挑战下智能体的韧性与体谅性参与
原标题:Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge
AI 摘要
该论文提出评估生成式AI智能体不能只看任务是否完成,还需考察「运营韧性」与「体谅性参与」两个维度。研究在医疗场景下模拟了120条轨迹,覆盖2个生成式AI模型和12项利益相关者任务,施加轻、中、重三级累积挑战,对比文本行动计划、提示式内部评估和结构化工作负荷与情感报告。结果显示,随挑战累积,智能体从自主恢复转向更多依赖人类,结构化报告中工作负荷与负面情绪上升但文本回应中很少表达压力;同时其适应方式从任务聚焦扩展到任务重构、关注他人、角色边界调整和更广泛协调。作者据此提出涉及坚持、注意力、角色边界、状态披露和升级的五项部署困境。
正文节选
Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge Abstract Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two