返回全部动态

Osprey:目标无关预训练打造更强推测解码草稿模型

原标题:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

arXiv cs.CL一手来源研究质量 90

AI 摘要

该论文提出 Osprey,一种目标无关的预训练方法,用于提升推测解码中草稿模型的泛化能力。现有草稿模型通常针对单一目标模型训练,工作负载变化时接受率会急剧下降。Osprey 利用现成预训练小语言模型,通过剪枝、目标无关的下一词预训练、词汇对齐、零初始化 QKV 扩展和蒸馏,实现跨目标模型迁移。实验显示,单个预训练 Osprey 骨干在 Qwen3-8B、Llama-3.3-70B-Instruct 和 229B MiniMax-M2.5 上均提升了平均接受长度,在域外和多语言数据上增益最大。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding Abstract Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pr


发布时间:2026-09-10 12:00
抓取时间:2026-09-10 12:08
来源机构:arXiv
阅读原文arxiv.org