返回全部动态

FreeToken:消费级硬件上的动态 MoE 推理引擎

原标题:FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution

InfoQ AI ML and Data Engineering研究质量 80

AI 摘要

UC Berkeley 和 MIT 的研究人员发布了 FreeToken,这是一个开源推理引擎,旨在通过动态协同执行,在消费级硬件上运行前沿 MoE 模型。FreeToken 采用 q* 策略动态分配 CPU 和 GPU 计算,并引入语义锚点检查点,显著提升解码和预填充速度。基准测试显示,FreeToken 在 RTX 4060 笔记本上运行 Qwen3.6-35B 达到约 39 tokens/s,并支持在单工作站 GPU 上运行 753B 的 GLM-5.2。社区反响热烈,认为这降低了自托管前沿模型的门槛,并减少对云 API 的依赖。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Researchers from UC Berkeley and MIT have introduced FreeToken, an open-source inference engine designed to bridge the gap between frontier Mixture-of-Experts (MoE) models and consumer-grade hardware. Co-authored by Databricks co-founders Matei Zaharia and Ion Stoica alongside Song Han, Kurt Keutzer and others, the project shifts the paradigm of edge AI from treating personal machines as constrained datacenter nodes to managing them as elastic, heterogeneous computing fabrics. While sparse MoE a


发布时间:2026-08-29 13:05
抓取时间:2026-08-29 13:09
来源机构:InfoQ
阅读原文infoq.com