Together AI 推出 Provisioned Throughput,提供开放模型预留推理容量
原标题:Open, convenient and predictable: Introducing Provisioned Throughput
AI 摘要
Together AI 推出了 Provisioned Throughput 服务,为前沿开放模型提供预留推理容量,采用基于 token 的定价和 99% 可用性 SLA。该服务允许客户购买 PTU 单位,以固定速率保证吞吐量,成本比 Claude Opus 4.8 低 90%。目前支持 MiniMax M3 和 GLM-5.2,覆盖北美和 EMEA 地区。此举旨在满足企业对可靠、可预测的开放模型推理需求,降低推理成本。
正文节选
We're excited to introduce Provisioned Throughput, reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA. Open weight models have become the essential ingredient for any company that wants an abundant AI future. But historically to use open weight models companies have had to choose between convenient but best-effort serverless or powerful but tunable dedicated inference. Provisioned Throughput is a new inference form factor that offers an alternate o