返回全部动态

DavidAU因基准问题封禁用户,分析揭示390个模型依赖过时基准

原标题:DAVIDAU DELETED MY POST AND BANNED ME FOR ASKING ABOUT BENCHMARKS — HERE IS THE FULL ANALYSIS (9B + 27B + 390 MODELS)

Hugging Face Blog一手来源观点质量 68

AI 摘要

用户因在DavidAU的模型页面询问现代基准测试分数而被删除帖子并封禁,随后发布了对DavidAU 390多个模型的分析。分析指出这些模型仅使用2018-2019年的饱和基准(如ARC-C),且声称的智能水平基于单一基准的小幅提升,并引用论文证明ARC-C的难度是评估方法造成的假象。该事件揭示了模型评估中基准选择的问题及社区管理中的争议。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

DAVIDAU DELETED MY POST AND BANNED ME FOR ASKING ABOUT BENCHMARKS — HERE IS THE FULL ANALYSIS (9B + 27B + 390 MODELS) I posted a discussion on the Qwen3.6-27B-Fable-Fusion-711 model page asking for modern benchmarks (SWE-bench, Terminal-Bench, GPQA, LiveCodeBench). DavidAU did not respond. He did not close the discussion. He deleted it and then blocked me from interacting with any of his repositories ("you're not enabled to interact with this user"). So now I'm posting this here. If he deletes t


发布时间:—
抓取时间:2026-08-12 19:45
来源机构:Hugging Face
阅读原文huggingface.co