返回全部动态

评估语义识别:基准效应与排名的稳定性分析

原标题:Stable Within, Unidentified Across: Semantic Identification of Benchmark Effects and Rankings

arXiv cs.SE一手来源研究质量 82

AI 摘要

该论文提出“评估语义识别”概念,研究评估结论是否在声明的语义族内保持不变。通过217条轨迹的实证分析,发现冻结的对比处理在特定语义族内产生零效应,而联合处理则无法识别效应,表明排名依赖规范。作者将分歧归因于单侧终端删除,并提供了可复用的族索引审计方法。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Stable Within, Unidentified Across: Semantic Identification of Benchmark Effects and Rankings Abstract Evaluation conclusions depend on evaluator-controlled semantics: legal references, scoreability, and aggregation. We call an artifact-defined endpoint evaluation-semantically identified when it is invariant over a declared family. A frozen 217-row analysis appears stable within its restricted contract family. In TraceElephant, yields precise task-disjoint adoption effects from to , whereas both


发布时间:2026-08-21 12:00
抓取时间:2026-08-21 14:24
来源机构:arXiv
阅读原文arxiv.org