返回全部动态
RubricForge:诱导无奖励评判标准以减少智能体评估中的过度信任
原标题:Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
AI 摘要
本文提出 RubricForge 方法,通过反射进化从少量带真实环境奖励标签的轨迹中自动生成评估智能体的评分标准,替代人工编写或微调评判模型。实验表明,该方法在保持与通用 G-Eval 评判器相当的整体一致性下,将虚假通过率降低约一半,从而减少对失败轨迹的过度信任,提升评估的可信度与可解释性。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation Abstract Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time. Such a judge is a reward-free proxy whose value depends on whether it can be trusted, yet existing judges either hand-write the scoring rubric, as in G-Eval, or fine-tune the judge’
发布时间:2026-08-17 12:00
抓取时间:2026-08-17 12:06
来源机构:arXiv