返回全部动态
Ollama v0.31.1 发布:Gemma 4 在 Apple Silicon 上提速近 90%
原标题:v0.31.1
AI 摘要
Ollama 发布 v0.31.1 版本,显著提升了 Gemma 4 在 Apple Silicon 上的运行速度,通过多 token 预测(MTP)技术,在编码代理基准测试中平均生成速度提升近 90%。该版本默认启用自动调优,无需配置,且不改变模型输出。此外,更新了 MLX 引擎和 llama.cpp 引擎,并改进了 Gemma 4 MoE 模型的加载性能。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
## Faster Gemma 4 on Apple Silicon <img width="1037" height="485" alt="Screenshot 2026-06-30 at 5 25 29 PM" src="https://github.com/user-attachments/assets/547d5076-090f-43c4-a661-938e11abc955" /> Gemma 4 is now significantly faster in Ollama on Apple Silicon, generating tokens nearly 90% faster on average across a coding-agent benchmark by leveraging multi-token prediction (MTP). Ollama auto-tunes how many tokens to draft as it runs, so the speedup is on by default, requires no configurat
发布时间:2026-07-01 06:10
抓取时间:2026-08-02 00:27
来源机构:Ollama