返回全部动态
可解释语言模型的规模化:可解释性随能力提升
原标题:Scaling Inherently Interpretable Language Models
AI 摘要
本研究挑战了可解释性以牺牲模型能力为代价的假设,提出将可解释性作为训练约束,与语言建模目标共同优化。实验表明,在自回归和扩散语言模型上,可解释性随计算规模提升而增强,模型表征变得更解耦且与人类可理解概念对齐。作者发布了Steerling-8B扩散语言模型,支持输出归因和概念引导干预,其性能可与使用2-16倍计算量训练的开放模型竞争。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Computer Science > Computation and Language Title:Scaling Inherently Interpretable Language Models View PDF Abstract:Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the languag
发布时间:2026-08-11 12:00
抓取时间:2026-08-11 12:14
来源机构:arXiv