返回全部动态
YCWTG 发布 gemma-4-31B-it 的 NVFP4A16 量化版
原标题:YCWTG/gemma-4-31B-it-NVFP4A16-GPTQ
AI 摘要
YCWTG 发布了基于 google/gemma-4-31B-it 的 NVFP4A16 量化版本模型,使用 llm-compressor 生成,模型大小从 62.6GB 降至 20.5GB,同时保持接近原版的推理性能(HLE 得分 18.1 vs 19.5)。该模型支持指令模式和思考模式,并针对视觉塔和 lm_head 层保留了 16 位精度以减少多模态性能损失。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
--- library_name: transformers base_model: - google/gemma-4-31B-it tags: - gemma4 - moe - NVFP4A16 - gptq - quantized - instruct license: apache-2.0 pipeline_tag: image-text-to-text --- <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/685e122d50df66f41587d406/XU7ovrDdgsNFAzahuyvI1.png" alt="Tone"> </p> Language [中文](https://huggingface.co/YCWTG/gemma-4-31B-it-NVFP4A16-GPTQ/blob/main/README_zh.md)|English ## Model Details This model is an **NVFP4
发布时间:2026-08-19 11:00
抓取时间:2026-08-19 09:25
来源机构:Hugging Face