返回全部动态

一致并非对齐:人类与LLM道德判断中的分歧基础

原标题:Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

arXiv cs.AI一手来源研究质量 84

AI 摘要

该研究通过一个包含500个条目的ETHICS衍生基准,比较了人类标注者和LLM在道德判断中的最终标签与支持理由,发现尽管标签一致性很高,但模型在道德理由上存在系统性分歧,如对伤害、尊重等类别的关注不同。作者认为,标签一致性不应等同于对齐,评估应关注模型表达的理由和原则。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments Abstract. Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 5


发布时间:2026-08-15 12:00
抓取时间:2026-08-15 12:00
来源机构:arXiv
阅读原文arxiv.org