返回全部动态
评估语义识别:基准效应与排名的稳定性分析
原标题:Stable Within, Unidentified Across: Semantic Identification of Benchmark Effects and Rankings
AI 摘要
该论文提出“评估语义识别”概念,研究评估结论是否在声明的语义族内保持不变。通过217条轨迹的实证分析,发现冻结的对比处理在特定语义族内产生零效应,而联合处理则无法识别效应,表明排名依赖规范。作者将分歧归因于单侧终端删除,并提供了可复用的族索引审计方法。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Stable Within, Unidentified Across: Semantic Identification of Benchmark Effects and Rankings Abstract Evaluation conclusions depend on evaluator-controlled semantics: legal references, scoreability, and aggregation. We call an artifact-defined endpoint evaluation-semantically identified when it is invariant over a declared family. A frozen 217-row analysis appears stable within its restricted contract family. In TraceElephant, yields precise task-disjoint adoption effects from to , whereas both
发布时间:2026-08-21 12:00
抓取时间:2026-08-21 14:24
来源机构:arXiv