返回全部动态

Together AI 基准测试:编码代理推理引擎性能领先

原标题:Benchmarking inference at scale: coding agents

Together AI Blog一手来源研究质量 84

AI 摘要

Together AI 发布了一项针对编码代理工作负载的大规模推理基准测试,结果显示其推理引擎在相同硬件上比最快的开源引擎吞吐量高 31%,且在饱和状态下 TTFT 快 2 倍。该基准测试模拟了长上下文、高并发的生产环境,并强调了 TTFT 和预填充压力等关键指标。性能提升源于全栈优化,包括 ThunderMLA 内核融合和自定义内核重写。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

On a production coding agent workload, Together Inference Engine delivers 31% more TPS than the next fastest OSS engine on the same hardware, and maintains 2× better TTFT at saturation. The gains come from full-stack optimization: ThunderMLA, custom kernel rewrites, and end-to-end profiling on real traffic. Most inference benchmarks measure a single user hitting a dedicated endpoint. The numbers look great. They're also useless for reasoning about production. In production, you're running dozens


发布时间:—
抓取时间:2026-08-03 01:12
来源机构:Together AI
阅读原文together.ai