返回全部动态

Primitive AI 发布 Qwen3.8-Flash-Next NVFP4 量化版,单 GPU 可运行

原标题:primitive-ai/Qwen3.8-Flash-Next-NVFP4

Hugging Face New and Trending Models一手来源模型发布质量 84

AI 摘要

Primitive AI 发布了 Qwen3.8-Flash-Next 的 NVFP4 量化版本,使 180B 参数的模型能在单张 96GB Blackwell GPU 上运行。该量化模型在 9 项基准测试中平均得分 92.2,支持 MTP 投机解码,并提供了详细的部署方案,包括 CPU 卸载和 NVMe 映射选项。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: other license_name: qwen-community-1.0 license_link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE base_model: Qwen/Qwen3.8-Flash-Next base_model_relation: quantized pipeline_tag: image-text-to-text library_name: transformers tags: - nvfp4 - quantized - vllm - modelopt - qwen3.8 - flash-next - single-gpu - speculative-decoding thumbnail: https://huggingface.co/primitive-ai/Qwen3.8-Flash-Next-NVFP4/resolve/main/assets/banner.png --- <p align="center"> <img src="


发布时间:2026-09-04 18:25
抓取时间:2026-09-04 18:26
来源机构:Hugging Face
阅读原文huggingface.co