返回全部动态

llama.cpp 发布 b10996:为 qwen3-coder 强制推理预算结束标签

原标题:b10996

llama.cpp Releases一手来源开源质量 58

AI 摘要

llama.cpp 发布 b10996 版本,主要变更是针对 qwen3-coder 在推理预算结束时强制输出 `\n</think>` 标签。该版本同时提供 macOS、iOS、Linux、Android、Windows 等多平台多后端的预编译二进制文件,包括 CUDA、Vulkan、ROCm、SYCL、OpenVINO 等加速方案。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

<details open> chat : force `\n</think>` on reasoning budget end for qwen3-coder (#28869) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/47858837> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10996/llama-b10996-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](https://github.com/ggml-org/llama.cpp/pull/23780) - [macOS Intel (x64)]


发布时间:2026-09-16 16:36
抓取时间:2026-09-16 17:16
来源机构:ggml-org
阅读原文github.com