返回全部动态

DeepSeek-V4.1-Flash 多精度 GGUF 量化版发布

原标题:smalinin/DeepSeek-V4.1-Flash-GGUF

Hugging Face New and Trending Models一手来源开源质量 79

AI 摘要

开发者 smalinin 在 Hugging Face 发布了 DeepSeek-V4.1-Flash 的 GGUF 量化版本,包含 Q2_K、IQ2_XXS、IQ3_XS、MXFP4 等多种混合精度方案,模型体积从 335GB 到 508GB 不等。这些文件需要配套的 llama.cpp 自定义分支 my_build_deepseek41 才能运行,且仅支持 CUDA。该架构包含稀疏注意力、超连接、MoE 层以及两个大型 n-gram Engram 查找表。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: mit base_model: deepseek-ai/DeepSeek-V4.1-Flash base_model_relation: quantized library_name: gguf pipeline_tag: image-text-to-text inference: false tags: - gguf - deepseek - deepseek-v4.1 - mixture-of-experts - llama.cpp - vision quantized_by: smalinin --- # DeepSeek-V4.1-Flash GGUF GGUF conversions and mixed-precision quantizations of [deepseek-ai/DeepSeek-V4.1-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) for the experimental DeepSeek-V4.1 runtime in [smalinin/l


发布时间:2026-09-24 09:01
抓取时间:2026-09-23 02:42
来源机构:Hugging Face
阅读原文huggingface.co