返回全部动态
Together AI 推出 LLM 推理自动扩缩容端点
原标题:Autoscaling endpoints for LLM inference
AI 摘要
Together AI 为其专用模型推理平台推出了基于推理引擎原生指标(如 in-flight 请求数、TTFT、GPU 利用率、token 吞吐量)的自动扩缩容功能。用户可设置副本范围、选择指标和目标值,并调整扩缩容窗口以平衡延迟和成本。文章通过实验对比了不同扩缩容策略,并提供了选择合适指标的建议。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
With Dedicated Model Inference on the Together AI platform you can get your deployments to autoscale on metrics the inference engine actually understands, such as in-flight requests, TTFT, GPU utilization, token throughput. You can set replica bounds, pick a metric and target, and then tune two windows that control how eagerly it scales up and how patiently it scales down. Understanding and choosing the right metric is important because it determines how your deployment will behave under peaky t
发布时间:—
抓取时间:2026-08-03 01:12
来源机构:Together AI