返回全部动态

Together AI 增强 GPU 集群:可靠性提升与运维控制

原标题:New in Together GPU Clusters: Reliability and control for production GPU clusters

Together AI Blog一手来源产品发布质量 82

AI 摘要

Together AI 发布了 GPU 集群的多项更新,包括被动健康检查、自动节点修复、Slinky 2.0 以及新的运维控制功能,旨在提升大规模训练和推理的可靠性。这些更新通过检测硬件故障、自动修复节点、增强 Slurm 稳定性,并提供 OIDC、启动脚本等自定义选项,帮助团队更高效地管理生产集群。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We’ve spent the last several weeks shipping a set of changes to Together GPU Clusters aimed at the operational reality of running training and inference at scale: hardware fails, schedulers leak, and teams outgrow the single-admin-kubeconfig workflow they started with. This post walks through what we shipped, why we built it the way we did, and what it means for the workloads you’re running on Together today. The changes group into two themes. The first is platform health: passive health checks,


发布时间:—
抓取时间:2026-08-03 01:12
来源机构:Together AI
阅读原文together.ai