返回全部动态

可解释语言模型的规模化:可解释性随能力提升

原标题:Scaling Inherently Interpretable Language Models

arXiv cs.CL一手来源研究质量 85

AI 摘要

本研究挑战了可解释性以牺牲模型能力为代价的假设,提出将可解释性作为训练约束,与语言建模目标共同优化。实验表明,在自回归和扩散语言模型上,可解释性随计算规模提升而增强,模型表征变得更解耦且与人类可理解概念对齐。作者发布了Steerling-8B扩散语言模型,支持输出归因和概念引导干预,其性能可与使用2-16倍计算量训练的开放模型竞争。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Computation and Language Title:Scaling Inherently Interpretable Language Models View PDF Abstract:Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the languag


发布时间:2026-08-11 12:00
抓取时间:2026-08-11 12:14
来源机构:arXiv
阅读原文arxiv.org