返回全部动态

ThunderAgent:实现 2 倍加速的智能体推理系统

原标题:ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

Together AI Blog一手来源研究质量 87

AI 摘要

Together AI 联合多所大学推出 ThunderAgent,一种面向智能体推理的高吞吐调度系统。它通过程序级调度缓解 KV 缓存抖动,在单节点上实现 2.5 倍吞吐提升,8 节点集群加速 2.4 倍,并保持近线性扩展。该系统已作为 Spotlight 论文被 ICML 2026 接收。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We introduce ThunderAgent, a system for high throughput agentic inference. By introducing a novel program abstraction for agentic LLM request scheduling, ThunderAgent achieves up to 2.5× higher single-node throughput in our synthetic data generation pipeline, and delivers 2.4× speedup on an 8-node cluster with near-linear throughput scaling with respect to GPU nodes. Key Results → More than 2× single-node throughput, with roughly 10× lower P50 latency at high concurrency → 2.4× speedup on 8 node


发布时间:—
抓取时间:2026-08-03 01:12
来源机构:Together AI
阅读原文together.ai