返回全部动态

可解释语言模型的规模化:Steerling-8B 实现训练内可解释性

原标题:Scaling Inherently Interpretable Language Models

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

Hugging Face 发布了一篇技术报告,提出将可解释性作为训练约束而非事后解释的方法。实验表明,在自回归和扩散语言模型中,可解释性随规模提升而增强。他们发布了 Steerling-8B 模型,该模型能归因输出到输入 token、概念和训练数据,支持闭环干预,且性能可与使用 2-16 倍计算量训练的模型竞争。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Scaling Inherently Interpretable Language Models Abstract Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on b


发布时间:
抓取时间:2026-08-11 13:52
来源机构:Hugging Face
阅读原文huggingface.co