返回全部动态

Agention 发布 Qwen3.8-Flash-Next ROCmFP4 量化版,可在 Strix Halo 上全 GPU 运行

原标题:agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF

Hugging Face New and Trending Models一手来源模型发布质量 85

AI 摘要

Agention 发布了 Qwen3.8-Flash-Next 的 ROCmFP4 量化版 GGUF 模型,该模型可在 128GB 统一内存设备(如 Strix Halo)的 GPU 上完全运行,占用 87.06 GiB,困惑度仅比未量化模型高 2.48%。该量化版本支持视觉和投机解码,并针对长上下文预填充进行了优化,在 128k 上下文下仍保持 138 t/s 的吞吐量。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- base_model: - Qwen/Qwen3.8-Flash-Next base_model_relation: quantized license: other license_name: qwen-community-1.0 license_link: LICENSE library_name: gguf pipeline_tag: text-generation tags: - gguf - rocmfp4 - rocmfpx - vulkan - strix-halo - qwen4exp - imatrix --- # Qwen3.8-Flash-Next ROCmFP4-FAST imatrix GGUF A 180 B model that runs **entirely on the GPU** of a 128 GB unified-memory box — 87.06 GiB at 4.23 bpw, within 2.5% perplexity of the unquantized model. Sized for the 96 GiB VRAM


发布时间:2026-09-10 07:05
抓取时间:2026-09-10 07:07
来源机构:Hugging Face
阅读原文huggingface.co