该对 LLM 相关性评判者礼貌吗?语气作为严厉度操作点偏移
原标题:Should I Be Polite to My LLM Relevance Judge? Tone as a Severity Operating-Point Shift
AI 摘要
该研究在 TREC DL19/DL20 上测试了提示语气对 LLM 相关性评判的影响,覆盖 8 个评判模型、5 个经分类器校准的礼貌等级及每级 3 种改写。结果显示语气效应强烈依赖模型,且主要体现为评判者「严厉度操作点」的偏移,而非判断能力提升;语气对基于校准的绝对指标(如 Cohen's κ)影响远大于对排序质量(NDCG@10)的影响。作者据此提出操作点解释,调和了此前关于礼貌提示利弊的矛盾发现,并指出语气是绝对相关性标签场景下的潜在效度威胁。
正文节选
Should I Be Polite to My LLM Relevance Judge? Tone as a Severity Operating-Point Shift Abstract. Large language models are increasingly used as relevance judges, yet their labels can shift with prompt surface form. We study one such feature—tone—on TREC DL19/DL20 query–passage pairs, across eight judge models, five classifier-calibrated politeness levels, and three paraphrases per level. Effects are strongly model-dependent: one judge shows a structured U-shaped response, whereas most show only