返回全部动态

Qwen3.8-27B 推出 16GB 显存适配的 GGUF 量化版

原标题:cHunter789/Qwen3.8-27B-i1-IQ4_KS_KT-GGUF

Hugging Face New and Trending Models一手来源模型发布质量 81

AI 摘要

cHunter789 在 Hugging Face 上发布了 Qwen3.8-27B 模型的 GGUF 量化版本,采用 ik_llama.cpp 项目的 IQ4_KS 和 IQ4_KT 量化技术,旨在让模型适配 16GB 显存的 NVIDIA 显卡。该量化版本通过 q4_0 KV 缓存量化支持最长 110k 上下文,并对比了与 mradermacher 量化版本的困惑度,结果显示其性能略优。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: apache-2.0 base_model: Qwen/Qwen3.6-27B tags: - gguf - image-text-to-text - qwen - ik_llama.cpp - nvidia - imatrix - 4-bit - iq4_ks - iq4_kt pipeline_tag: image-text-to-text quantized_by: cHunter789 --- # Qwen3.8-27B-i1-IQ4_KS_KT-GGUF **This quantization was created to allow the entire model to fit into the memory of an NVIDIA graphics card with 16GB of VRAM.** **Importan - please use "export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1"** ``` When GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 is def


发布时间:2026-08-22 21:40
抓取时间:2026-08-22 21:40
来源机构:Hugging Face
阅读原文huggingface.co