混合线性注意力大语言模型中的大规模激活:预注意力尖峰与尖峰间平台
原标题:Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
AI 摘要
本研究首次系统性地探讨了混合线性注意力大语言模型中的大规模激活现象,揭示了两种与架构对齐的形态:全注意力层前的预注意力尖峰和线性注意力层间的尖峰间平台。研究发现,随着全注意力密度增加,这些形态逐渐连接并恢复为全注意力模型的稳定形态。该现象在多种架构、配置和数据域中普遍存在,且受输出门控机制的不对称影响。
正文节选
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Abstract Massive activations in hybrid linear-attention LLMs exhibit pre-attention spikes and inter-spike plateaus governed by cancellation timing, with morphology recovering at full-attention limits. We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately