IsValorum 发布 Qwen3.8-35B MoE 定制 GGUF 量化版
原标题:IsValorum/Qwen3.8-35B-A3B-Distill-MTP-APEX-I-MiniPlus-V2.1-GGUF
AI 摘要
IsValorum 在 Hugging Face 发布了 Qwen3.8-35B-A3B-Distill-MTP-APEX-I-MiniPlus-V2.1 的 GGUF 量化版本,基于 empero-ai/Qwen3.8-35B-A3B-Distill 模型。该版本针对 13-14GB 显存/内存上限优化,采用逐张量定制量化策略,保留 F32 路由门、Q6_K 输出头等关键组件,在 WikiText-2 上困惑度约 5.3952,接近 Q5_K_L 质量层级。作者同时提示上游 BF16 模型存在长文本重复问题,并提供了无审查的 Abliterated 姊妹版本。
正文节选
--- base_model: empero-ai/Qwen3.8-35B-A3B-Distill quantized_by: IsValorum library_name: gguf language: - en - zh - es - fr - de - pt - it - ru - ja - ko - vi - th - ar tags: - gguf - quantized - quantization - apex - apex-quant - custom-quantization - unsloth-studio - moe - reasoning - llama.cpp - qwen35moe license: apache-2.0 pipeline_tag: text-generation --- # Qwen3.8-35B-A3B-Distill-MTP APEX-I-MiniPlus-V2.1 GGUF ### *The Definitive Frontier MoE · Efficient System RAM Offload · Full 256K Cont