返回全部动态

Minions:端侧小模型与云端大模型协作降低 LLM 成本

原标题:Summary

Together AI Blog一手来源研究质量 85

AI 摘要

Together AI 的研究团队提出了一种名为 Minions 的方法,通过让小型端侧模型与云端前沿模型协作,将大量 LLM 工作负载转移到消费设备上,仅本地读取长上下文,从而降低云成本且几乎不损失质量。实验表明,Minions 在保持 97.9% 准确率的同时,成本仅为云端方案的 17.5%。该方法通过分解-执行-聚合循环,解决了小模型长上下文和多步指令的弱点,并展示了模型选择、推理计算扩展和顺序通信对成本-准确率权衡的影响。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Summary We present Minions, a method that shifts a substantial portion of LLM workloads to consumer devices by having small on-device models collaborate with frontier models in the cloud. By only reading long contexts locally, we reduce cloud costs with minimal or no quality degradation. We imagine a future where an “intelligence layer” running persistently on-device interacts with frontier models in the cloud to deliver applications with cost-effective, “always-on” intelligence. The Tailwinds:


发布时间:—
抓取时间:2026-08-04 02:56
来源机构:Together AI
阅读原文together.ai