超越对错:评估大语言模型中的二阶社会推理
原标题:Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
AI 摘要
该研究提出了一种评估大语言模型(LLM)二阶社会推理(元规范)的新框架,涵盖情绪评估和行为反应两个维度,并发布了包含450个规范违反场景的多视角数据集NormReact。研究发现,当前LLM在预测社会制裁时比人类更严厉,过度预测负面制裁,且随社会距离增加与人类判断的一致性下降。这表明AI系统在规范敏感领域可能产生扭曲的社会调节图景,过度代表惩罚而低估现实中的宽容与关系校准。
正文节选
Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models Abstract Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., ‘do not steal’). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people