Qwen3.6-27B PrismaSCOUT 量化版发布:NVFP4+BF16 混合精度,体积缩小 11%
原标题:rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm
AI 摘要
Hugging Face 上发布了 Qwen3.6-27B 的 PrismaSCOUT 量化版本,采用 NVFP4+BF16 混合精度,大小为 20.17 GB,比之前的 5.5 bpp 版本小 11%。该版本通过 PrismaSCOUT 算法基于端到端 KL 散度选择位分配,而非逐层代理指标,旨在减少量化漂移。实测平均 KL 为 0.0779,相比其 BF16 源模型,但低于 PrismaScout-AQUA 的 0.0402。该模型适用于 vLLM 在 NVIDIA Blackwell 上部署,可节省显存以支持更长上下文和更高并发。
正文节选
--- license: apache-2.0 base_model: Qwen/Qwen3.6-27B library_name: vllm tags: - qwen3_5 - compressed-tensors - nvfp4 - vllm - speculative-decoding - prismaquant - prismascout - mixed-precision - blackwell --- # Qwen3.6-27B — PrismaSCOUT (Blackwell, NVFP4 + BF16) **Newest version:** [PrismaScout-AQUA](https://huggingface.co/rdtand/PrismaScout-AQUA) is the current 20 GB text-only successor. This Qwen3.6 release remains available, and you are welcome to continue using it if you