返回全部动态

FP8 KV-Cache 在 Intel Arc Pro B70 上实现 2 倍容量与长上下文吞吐增益

原标题:FP8 KV-Cache on Intel® Arc™ Pro B70: 2× Capacity with strong Long Context Throughput Gains

Hugging Face Blog一手来源研究质量 84

AI 摘要

Hugging Face 博客评估了 Intel Arc Pro B70 上 FP8 KV-cache 量化在 vLLM 中的表现,显示相比 BF16,FP8 KV 在所有测试模型上实现了确定性 2 倍 KV 缓存容量提升,并在长上下文(16K-32K)场景下带来最高 42.3% 的吞吐量增益。FP8 KV 在 10 个模型中的 8 个上接近无损,但 DeepSeek-R1-Distill-Qwen-7B 和 Gemma-3-1B-IT 需要模型特定验证。该优化通过单一运行时标志启用,无需重训练或架构更改,但应作为选择性优化而非通用默认设置。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

- Capacity: FP8 KV delivers a deterministic 2.0× KV-cache token-capacity increase over BF16 across all ten evaluated models (1B–72B) and TP configurations (TP=1, 2, 4) - Concurrency: This directly translates into about 2.0× higher raw 4K session capacity at the same max_model_len. For example, Qwen2.5-14B-Instruct (TP=2) scales from 48.52 → 97.05 sessions, and Mistral-Small-24B-Instruct-2501 (TP=2) from 43.27 → 86.53 sessions, making FP8 KV a strong enabler for capacity-constrained deployments.


发布时间:—
抓取时间:2026-08-12 16:40
来源机构:Hugging Face
阅读原文huggingface.co