返回全部动态
Ollama v0.32.4 发布:支持 Apple GPU 并优化 Qwen3 MoE 性能
原标题:v0.32.4
AI 摘要
Ollama 发布 v0.32.4 版本,新增对 Apple GPU 上 Laguna 模型的支持(通过 MLX 引擎),并优化了推测解码草稿模型的量化处理。同时修复了 Qwen3 MoE 模型中不同量化专家解码的问题,并提升了打包门控/上投影的速度(在 M5 Max 上提升约 4-9%)。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
## What's Changed - Support Laguna on Apple GPUs via the MLX engine - Quantize draft-model output heads at the requested type when creating speculative-decoding drafts. - Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max). **Full Changelog**: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4
发布时间:2026-07-25 10:22
抓取时间:2026-08-02 00:27
来源机构:Ollama