返回全部动态

GLM-5.3-Flash NVFP4 量化版发布:体积缩减70%

原标题:LibertAIDAI/GLM-5.3-Flash-NVFP4

Hugging Face New and Trending Models一手来源模型发布质量 87

AI 摘要

LibertAI 发布了 GLM-5.3-Flash 的 NVFP4 量化版本,将模型从 598.5 GiB 压缩至 181 GiB,体积减少 70%,余弦相似度达 0.99665。该量化仅针对路由专家 FFN,保留注意力、视觉塔等关键部分为 BF16,确保多模态行为与原始模型一致。模型支持 vLLM 和 SGLang,但需要特定配置,且已修复 GB10 上的兼容性问题。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: mit base_model: zai-org/GLM-5.3-Flash base_model_relation: quantized quantized_by: LibertAIDAI tags: - nvfp4 - blackwell - sglang - vllm - glm - glm-5 - glm5_next - moe - multimodal - modelopt language: - en - zh pipeline_tag: image-text-to-text --- <div align="center"> # GLM-5.3-Flash · NVFP4 ### 320B total · 18B active · natively multimodal · 1M context **598.5 GiB → 181 GiB** &nbsp;·&nbsp; **−70%** &nbsp;·&nbsp; round-trip cosine **0.99665** [![Base](


发布时间:2026-08-30 20:44
抓取时间:2026-08-30 20:46
来源机构:Hugging Face
阅读原文huggingface.co