mlx-optiq 发布 Gemma-4 12B QAT 混合精度 4-bit 量化模型
原标题:mlx-community/gemma-4-12B-it-qat-OptiQ-4bit
AI 摘要
mlx-optiq 发布了 mlx-community/gemma-4-12B-it-qat-OptiQ-4bit,这是基于 Google 量化感知训练(QAT)版 Gemma-4 12B 的 4-bit 混合精度 MLX 量化模型。它采用敏感度引导的逐层比特分配(157 个组件用 8-bit,171 个用 4-bit,平均 5.25 bits-per-weight),在六项基准上的 Capability Score 比同基座的均匀 4-bit 量化高 1.37 分。模型面向 Apple Silicon 本地推理,支持图像+文本输入和投机解码草稿模型。
正文节选
--- library_name: mlx license: apache-2.0 license_link: https://ai.google.dev/gemma/docs/gemma_4_license pipeline_tag: text-generation base_model: google/gemma-4-12B-it-qat-q4_0-unquantized tags: - mlx - quantized - mixed-precision - 4bit - 8bit - optiq - qat - apple-silicon - text-generation - image-text-to-text - gemma-4 --- # mlx-community/gemma-4-12B-it-qat-OptiQ-4bit > **Built with [mlx-optiq](https://mlx-optiq.com)**, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally