返回全部动态

Qwen3.8-27B-heretic-dflash:非官方 5 层 DFlash 投机解码草稿模型

原标题:jfan/Qwen3.8-27B-heretic-dflash

Hugging Face New and Trending Models一手来源模型发布质量 83

AI 摘要

jfan 发布了非官方的 Qwen3.8-27B-heretic-dflash 模型,这是一个基于 DFlash 架构的 5 层投机解码草稿模型,专为目标模型 trohrbaugh/Qwen3.8-27B-heretic-ara 设计。该模型目前训练至 10k/60k 步,在 Apple M5 Pro 上实现了约 41.6% 的解码加速,支持 MLX、vLLM 和 SGLang 等推理框架。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: apache-2.0 base_model: trohrbaugh/Qwen3.8-27B-heretic-ara tags: - dflash - speculative-decoding - fast-inference - mlx - vllm - sglang - qwen3 - code - function-calling pipeline_tag: text-generation --- # Qwen3.8-27B-heretic-dflash: Unofficial 5-Layer DFlash Speculative Drafter *WARNING: This model is still in training. Currently completed 10k/60k steps.* **Note: This is not an official drafter, but is instead trained by continuing from [`z-lab/Qwen3.6-27B-DFlash`](https://huggin


发布时间:2026-08-25 06:51
抓取时间:2026-08-20 22:23
来源机构:Hugging Face
阅读原文huggingface.co