返回全部动态
Qwen3.8-27B 推出 W8A8 INT8 量化版,性能超越官方 FP8
原标题:Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8
AI 摘要
Freaksterz 发布了 Qwen3.8-27B 的 W8A8 INT8 量化版本,采用 SmoothQuant 和激活感知 GPTQ 技术,在保持低 KL 散度的同时实现比 W8A16 更快的预填充速度。该模型兼容 DFlash2 投机解码,并提供了详细的评估数据和部署指南。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
--- license: apache-2.0 base_model: Qwen/Qwen3.8-27B tags: [qwen3.8, w8a8, int8, smoothquant, gptq, compressed-tensors, vllm, dflash2, ampere] --- # Qwen3.8-27B — SmoothQuant + activation-aware GPTQ, W8A8 INT8 (v3) **v3 (`main`, 2026-09-03):** mean full-vocab KLD vs BF16 **0.00556** — lower than the official `Qwen/Qwen3.8-27B-FP8` on the same harness (0.00584) — and the stock **DFlash2 drafter works unchanged** (acceptance length 4.34, identical to a W8A16 target). No residual-stream rotation.
发布时间:2026-09-04 14:35
抓取时间:2026-09-04 14:37
来源机构:Hugging Face