返回全部动态

Ollama v0.31.1 发布:Gemma 4 在 Apple Silicon 上提速近 90%

原标题:v0.31.1

Ollama Releases一手来源产品发布质量 78

AI 摘要

Ollama 发布 v0.31.1 版本,显著提升了 Gemma 4 在 Apple Silicon 上的运行速度,通过多 token 预测(MTP)技术,在编码代理基准测试中平均生成速度提升近 90%。该版本默认启用自动调优,无需配置,且不改变模型输出。此外,更新了 MLX 引擎和 llama.cpp 引擎,并改进了 Gemma 4 MoE 模型的加载性能。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

## Faster Gemma 4 on Apple Silicon <img width="1037" height="485" alt="Screenshot 2026-06-30 at 5 25 29 PM" src="https://github.com/user-attachments/assets/547d5076-090f-43c4-a661-938e11abc955" /> Gemma 4 is now significantly faster in Ollama on Apple Silicon, generating tokens nearly 90% faster on average across a coding-agent benchmark by leveraging multi-token prediction (MTP). Ollama auto-tunes how many tokens to draft as it runs, so the speedup is on by default, requires no configurat


发布时间:2026-07-01 06:10
抓取时间:2026-08-02 00:27
来源机构:Ollama
阅读原文github.com