返回全部动态

Rasch测量理论用于LLM评估:以LLM作为评分者的案例研究

原标题:Rating the Raters: Rasch Measurement Theory for LLM Evaluation

arXiv cs.AI一手来源研究质量 87

AI 摘要

该研究将Rasch测量理论应用于LLM评估,以LLM作为评分者标注仇恨言论为例,分析了九个LLM与人类评分者的差异。研究发现LLM在严重性、项目校准、问题顺序鲁棒性、目标身份敏感性和评分量表使用上系统性地偏离人类,这些差异在标准评估中会被掩盖。作者主张RMT应纳入LLM评估工具箱。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Rating the Raters: Rasch Measurement Theory for LLM Evaluation Abstract LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our un


发布时间:2026-09-01 12:00
抓取时间:2026-08-31 12:11
来源机构:arXiv
阅读原文arxiv.org