返回全部动态
Together AI 基准测试:编码代理推理引擎性能领先
原标题:Benchmarking inference at scale: coding agents
AI 摘要
Together AI 发布了一项针对编码代理工作负载的大规模推理基准测试,结果显示其推理引擎在相同硬件上比最快的开源引擎吞吐量高 31%,且在饱和状态下 TTFT 快 2 倍。该基准测试模拟了长上下文、高并发的生产环境,并强调了 TTFT 和预填充压力等关键指标。性能提升源于全栈优化,包括 ThunderMLA 内核融合和自定义内核重写。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
On a production coding agent workload, Together Inference Engine delivers 31% more TPS than the next fastest OSS engine on the same hardware, and maintains 2× better TTFT at saturation. The gains come from full-stack optimization: ThunderMLA, custom kernel rewrites, and end-to-end profiling on real traffic. Most inference benchmarks measure a single user hitting a dedicated endpoint. The numbers look great. They're also useless for reasoning about production. In production, you're running dozens
发布时间:—
抓取时间:2026-08-03 01:12
来源机构:Together AI