返回全部动态
Qwen3.8-27B EXL3 量化版:更小下载、更快启动、更高保真
原标题:malaiwah/Qwen3.8-27B-EXL3-K5K6-hydrated
AI 摘要
malaiwah 发布了 Qwen3.8-27B 的 EXL3 量化版本,其中注意力权重在磁盘上以 K6 精度序列化,相比 BF16 注意力版本,下载体积减少 29%,冷启动速度提升 5.4 倍,且与 BF16 的 KL 散度更低。该模型需要特定的 Gilded Gnosis vLLM 分支才能运行,属于实验性研究产物。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
--- license: apache-2.0 base_model: Qwen/Qwen3.8-27B base_model_relation: quantized pipeline_tag: image-text-to-text library_name: vllm tags: - exl3 - exllamav3 - trellis - mixed-precision - quantized - qwen3.8 - vision-language - gilded-gnosis --- # Qwen3.8-27B EXL3 K5/K6, attention serialized on disk — 9 GB smaller, 5x faster cold start, and measurably closer to BF16 > **Requires a custom runtime.** Does **not** load in upstream vLLM, SGLang, TensorRT-LLM, > llama.cpp, transf
发布时间:2026-08-19 19:46
抓取时间:2026-08-17 21:17
来源机构:Hugging Face