GLM-5.3-Flash MLX 4-bit 量化版发布,支持 Apple Silicon 与原生 MTP
原标题:Vontra/GLM-5.3-Flash-MLX-4bit-MTP
AI 摘要
Vontra 发布了基于 zai-org/GLM-5.3-Flash 的 MLX 4-bit 量化版本,支持 Apple Silicon,并保留了原生 MTP(多 token 预测)层。该模型采用 4-bit 仿射量化,组大小为 64,总下载大小约 181.7 GB,支持 1M 上下文,架构为 glm5_next 多模态稀疏 MoE。在 Apple M3 Ultra 上,基线生成速度约为 6.27 tokens/s,但原生 MTP 需要特定运行时支持。
正文节选
--- library_name: mlx license: mit license_link: LICENSE base_model: zai-org/GLM-5.3-Flash base_model_relation: quantized pipeline_tag: image-text-to-text language: - en - zh tags: - mlx - mlx-vlm - omlx - glm - glm5 - glm5-next - native-mtp - speculative-decoding - mixture-of-experts - multimodal - vision-language - quantized - apple-silicon --- <div align="center"> <a href="https://huggingface.co/zai-org/GLM-5.3-Flash"> <img src="zai-logo.png" width="96" h