DeepSeek-V4-Flash 定制 GGUF 发布,面向 AMD Strix Halo 优化
原标题:otheru/DeepSeek-V4-Flash-Strix-Halo-GGUF
AI 摘要
Hugging Face 上发布了 DeepSeek-V4-Flash-0731 的定制 GGUF 量化版本,专为 AMD Strix Halo 平台的 Ember 运行时设计,采用 ROCmFPx 自定义张量类型,不兼容主流 llama.cpp 等运行时。该版本经过 abliteration 和重要性矩阵校准,并附带 DSpark 草稿模型,在 AMD Ryzen AI Max+ 395 上测得 34.16 tok/s 的中位解码速度,但量化质量尚未评估。
正文节选
--- license: other license_name: deepseek base_model: deepseek-ai/DeepSeek-V4-Flash-0731 tags: - gguf - rocmfp - rocmfpx - strix-halo - gfx1151 - amd - deepseek-v4 - deepseek-v4-0731 - moe - imatrix - abliterated pipeline_tag: text-generation --- # DeepSeek-V4-Flash-0731 — Strix Halo ROCmFPx GGUF > [!WARNING] > This is a custom GGUF for **[Ember](https://github.com/otheru-ai/ember)**, an > ROCmFPx-aware DeepSeek-V4 runtime for AMD Strix Halo (`gfx1151`). It uses > custom tensor types that main