识别、模拟与拒绝:LLM 智能体经典心理效应的污染感知研究
原标题:Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
AI 摘要
研究者提出 PsyAgentBench 基准,在 LLM 智能体上重跑经典心理学实验,通过「命名/盲测」与「教科书/反事实」交叉设计区分模型是真正具有类人偏差、识别实验后表演、还是复现记忆中的响应统计。在五个范式、最多三个开源模型家族、41,904 次试验中,类人效应通过不同路径出现,且一句人格提示可消除、削弱或反转这些效应。作者据此主张标量偏差易感性评分掩盖了结构差异,应改用复制剖面报告。
正文节选
[Page 1] Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents Joy Bose Independent Researcher, Bengaluru, India joy.bose@ieee.org Abstract. An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark