GLM-5.3-Flash W4A16 INT4 量化版发布,体积减70%
原标题:canada-quant/GLM-5.3-Flash-W4A16-MTP
AI 摘要
canada-quant 发布了 GLM-5.3-Flash 的 W4A16(INT4)量化版本,仅对 36,288 个路由专家 GEMM 做 GPTQ 对称 INT4 量化,注意力、路由、共享专家、嵌入、视觉塔和 MTP 头保持 BF16。模型体积从约 599 GiB 降至 177.7 GiB(-70%),可在 H100、H200、RTX PRO 6000 和 DGX Spark 上运行,2×DGX Spark 与 2×H200 配置支持完整 1,048,576 token 上下文。质量与 NVFP4 参考在 AIME 2025、GSM8K、GPQA-Diamond 上处于噪声范围内,吞吐在 H100 和 RTX PRO 6000 上持平或更优。
正文节选
--- language: - en - zh license: mit library_name: transformers pipeline_tag: image-text-to-text base_model: zai-org/GLM-5.3-Flash base_model_relation: quantized tags: - glm-5.3 - glm5_next - quantized - w4a16 - int4 - gptq - compressed-tensors - mtp - speculative-decoding - dflash2 - vision - multimodal - vllm - dgx-spark - 1m-context - blackwell - hopper --- # GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP INT4 weight-only quantization of [zai-org/GLM-5.3-Flash