返回全部动态

自然语言数学证明的低成本自动评判方法

原标题:Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

arXiv cs.CL一手来源研究质量 84

AI 摘要

该研究探讨了使用廉价开源权重模型作为自然语言数学证明的自动评判者。在IMO-GradingBench的200个实例验证样本上,三个廉价评判模型(GPT-OSS 120B、DeepSeek-V4 Flash、Gemma-4 31B)与人类通过/失败判断的一致性在统计上与Claude Opus 4.7和Gemini 3.1 Pro等前沿模型相当,但成本低至100倍。扩展到完整1000实例基准后,要求一致通过(all-three-pass)的共识规则达到最高的通过一致性和精确度。主要发现是廉价评判模型在成本低一到两个数量级的情况下与前沿模型竞争力相当,建议默认使用all-three-pass规则,但需独立复制验证。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Computation and Language Title:Cost-Effective Automated Judging of Natural-Language Mathematical Proofs View PDF HTML (experimental) Abstract:Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IM


发布时间:2026-08-04 12:00
抓取时间:2026-08-04 12:02
来源机构:arXiv
阅读原文arxiv.org