DavidAU因基准问题封禁用户,分析揭示390个模型依赖过时基准
原标题:DAVIDAU DELETED MY POST AND BANNED ME FOR ASKING ABOUT BENCHMARKS — HERE IS THE FULL ANALYSIS (9B + 27B + 390 MODELS)
AI 摘要
用户因在DavidAU的模型页面询问现代基准测试分数而被删除帖子并封禁,随后发布了对DavidAU 390多个模型的分析。分析指出这些模型仅使用2018-2019年的饱和基准(如ARC-C),且声称的智能水平基于单一基准的小幅提升,并引用论文证明ARC-C的难度是评估方法造成的假象。该事件揭示了模型评估中基准选择的问题及社区管理中的争议。
正文节选
DAVIDAU DELETED MY POST AND BANNED ME FOR ASKING ABOUT BENCHMARKS — HERE IS THE FULL ANALYSIS (9B + 27B + 390 MODELS) I posted a discussion on the Qwen3.6-27B-Fable-Fusion-711 model page asking for modern benchmarks (SWE-bench, Terminal-Bench, GPQA, LiveCodeBench). DavidAU did not respond. He did not close the discussion. He deleted it and then blocked me from interacting with any of his repositories ("you're not enabled to interact with this user"). So now I'm posting this here. If he deletes t