DeepSeek-V4-Flash 0731 推出 92GB GGUF 量化版,支持 GB10 全量运行
原标题:twaggs88/DeepSeek-V4-Flash-REAP25-DSpark-ds4-GGUF
AI 摘要
twaggs88 发布了 DeepSeek-V4-Flash-0731 的 GGUF 量化版本,这是一个 284B 参数的 MoE 模型,通过 IQ2_XXS_MMQ、MXFP4 和 MXFP8 等量化技术压缩至 92GB,可在 NVIDIA GB10 设备上全量运行。该版本包含完整的 256 个专家,并集成了 DSpark 投机解码草稿模型,支持长上下文和并发会话。文件仅兼容 pulsar 引擎,不适用于 llama.cpp 等标准加载器。
正文节选
--- license: mit base_model: deepseek-ai/DeepSeek-V4-Flash-DSpark-0731 pipeline_tag: text-generation tags: - gguf - deepseek - moe - quantized - 2-bit - mxfp4 - speculative-decoding - gb10 --- # DeepSeek-V4-Flash 0731 — IQ2_XXS_MMQ · MXFP4 · MXFP8 · DSpark (ds4 GGUF) A 92 GB single-file build of [DeepSeek-V4-Flash-0731](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark-0731) (284B MoE, 13B active, 1M context) sized to run **fully resident on one NVIDIA GB10 (128 GB un