返回全部动态

llama.cpp 发布 b11118:新增 hex-dma 直接映射 DMA 缓存

原标题:b11118

llama.cpp Releases一手来源开源质量 61

AI 摘要

llama.cpp 发布 b11118 版本,核心变更是引入 hex-dma 直接映射 DMA 缓存,用于更好地处理 HVX 上的 flash attention 掩码(FA mask)。该版本同时提供 macOS/iOS、Linux、Android、Windows 等多平台预编译包,覆盖 CPU、CUDA、Vulkan、ROCm、SYCL、OpenVINO 等后端。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

<details open> hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling (#29282) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/49404271> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b11118/llama-b11118-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](https://github.com/ggml-org/llama.cpp/pull/2378


发布时间:2026-09-23 10:10
抓取时间:2026-09-23 10:52
来源机构:ggml-org
阅读原文github.com