返回全部动态

llama.cpp b10255 扩展 SYCL SDPA 支持非 FP16 KV 缓存

原标题:b10255

llama.cpp Releases一手来源产品发布质量 83

AI 摘要

llama.cpp 发布 b10255 版本,扩展了 SYCL oneDNN SDPA 路径以支持非 FP16 的 KV 缓存(包括 Q4_0-Q8_0 和 FP32),通过在设备上将 K/V 反量化为 FP16 后送入 SDPA 图,使融合的 systolic 内核与原生 FP16 路径一致运行。该版本还包含流同步修复和移除 V_is_K_view 别名,并提供了多平台二进制下载。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

<details open> Extended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0–Q8_0 and FP32) (#25874) * sycl: extend oneDNN SDPA to Q4_0-Q8_0 and F32 KV caches Extends the oneDNN SDPA path (PR #25222) to handle non-F16 KV caches by dequantizing or converting K/V to dense FP16 on-device before feeding them into the SDPA graph. The fused systolic kernel then runs identically to the native FP16 path. Supported KV types: - Q4_0, Q4_1, Q5_0, Q5_1, Q8_0: to_fp16_sycl / to_fp16_nc_sycl - F32: cont_to_f1


发布时间:2026-08-04 13:39
抓取时间:2026-08-04 14:04
来源机构:ggml-org
阅读原文github.com