返回全部动态

GLM-5.3-Flash W4A16 INT4 量化版发布,体积减70%

原标题:canada-quant/GLM-5.3-Flash-W4A16-MTP

Hugging Face New and Trending Models一手来源模型发布质量 80

AI 摘要

canada-quant 发布了 GLM-5.3-Flash 的 W4A16(INT4)量化版本,仅对 36,288 个路由专家 GEMM 做 GPTQ 对称 INT4 量化,注意力、路由、共享专家、嵌入、视觉塔和 MTP 头保持 BF16。模型体积从约 599 GiB 降至 177.7 GiB(-70%),可在 H100、H200、RTX PRO 6000 和 DGX Spark 上运行,2×DGX Spark 与 2×H200 配置支持完整 1,048,576 token 上下文。质量与 NVFP4 参考在 AIME 2025、GSM8K、GPQA-Diamond 上处于噪声范围内,吞吐在 H100 和 RTX PRO 6000 上持平或更优。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- language: - en - zh license: mit library_name: transformers pipeline_tag: image-text-to-text base_model: zai-org/GLM-5.3-Flash base_model_relation: quantized tags: - glm-5.3 - glm5_next - quantized - w4a16 - int4 - gptq - compressed-tensors - mtp - speculative-decoding - dflash2 - vision - multimodal - vllm - dgx-spark - 1m-context - blackwell - hopper --- # GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP INT4 weight-only quantization of [zai-org/GLM-5.3-Flash


发布时间:2026-09-25 19:58
抓取时间:2026-09-25 20:00
来源机构:Hugging Face
阅读原文huggingface.co