返回全部动态

Qwen3.8-Flash-Next 推出 INT4 量化版,支持 Ampere GPU 运行

原标题:VnimanieAI/Qwen3.8-Flash-Next-W4A16

Hugging Face New and Trending Models一手来源模型发布质量 88

AI 摘要

VnimanieAI 发布了 Qwen3.8-Flash-Next 的 INT4 量化版本,这是该模型首个公开的 INT4 量化,可在消费级 Ampere GPU(如 RTX 3090)上运行,而官方 FP8 版本不支持这些 GPU。量化采用 W4A16 格式,模型大小从 335GB 降至 168GB,通过 vLLM 部署,支持专家并行和 MTP 推测解码。该版本针对 Ampere GPU 优化,解决了 FP8 和 NVFP4 在旧硬件上的兼容性问题。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: other license_name: qwen-community-1.0 license_link: LICENSE base_model: Qwen/Qwen3.8-Flash-Next base_model_relation: quantized pipeline_tag: image-text-to-text library_name: transformers tags: - qwen4_exp - compressed-tensors - w4a16 - int4 - 4-bit - vllm - ampere - rtx-3090 - conversational --- # Qwen3.8-Flash-Next-W4A16 INT4 (W4A16, group-128, symmetric) quantization of [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) in `compressed-te


发布时间:2026-09-02 20:50
抓取时间:2026-09-02 20:52
来源机构:Hugging Face
阅读原文huggingface.co