返回全部动态

Hugging Face 药物基准揭示数据分割与标签噪声对评估的关键影响

原标题:We changed one line and the benchmark score moved 0.21 AUROC

Hugging Face Blog一手来源研究质量 83

AI 摘要

Hugging Face 团队在构建药物预测基准 LEADBOARD 时发现,数据分割方式对模型性能评估影响巨大:时间分割与随机分割在同一数据集上 AUROC 相差 0.21。此外,跨文献的标签噪声(噪声下限)显著,hERG 数据中 90% 的重复测量差异超过 20 倍。团队因此采用时间或骨架分割,并公开噪声下限以帮助解读排名。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

LEADBOARD - ADMET, Kinase and Toxicity Prediction Benchmark Benchmark for drug prediction tools - ADMET, kinase, tox This post is mostly about two numbers we ran into while building it, because they changed what we thought the thing should be. hERG was the board we built first. It's the potassium channel that, when a drug blocks it, gives you a QT interval problem and a dead clinical program. Everybody screens for it early, so there's a lot of public data, which made it a good place to shake out


发布时间:—
抓取时间:2026-08-24 15:42
来源机构:Hugging Face
阅读原文huggingface.co