Together AI 定制投机解码加速 DeepSeek-R1 推理
原标题:Boosting DeepSeek-R1’s Speed with Customized Speculative Decoding
AI 摘要
Together AI 在博客中展示了为 DeepSeek-R1 定制投机解码(speculative decoding)加速器的方法,通过基于客户工作负载微调,相比其基础加速器可实现 1.23-1.45 倍解码速度提升和约 25% 成本降低,相比传统逐 token 预测则达到 1.85-2.97 倍加速和约 55% 成本削减。该方法利用客户专用端点的数据分布,训练定制加速器,并随数据量增加(如 20M 或 50M tokens)进一步提升性能。
正文节选
TLDR: In this blog post, we show that using a custom speculator—trained on your own Deepseek-R1 inference traffic—can yield 1.23-1.45x speedups during decoding (tokens/second), and ~25% reduction in overall cost (same throughput with fewer GPU-hours), relative to Together’s state-of-the-art base speculator. This translates to 1.85-2.97x speedup and ~55% cost reductions when compared to conventional next token prediction. Please reach out to our sales team to learn how to get started with a custo