NVIDIA NVLink 6 为 AI 工厂提供多层弹性
原标题:How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories
AI 摘要
NVIDIA 在技术博客中介绍 NVLink 6 面向大规模 AI 工厂的多层弹性架构,覆盖物理层、链路层、系统冗余和应用软件层。物理层通过轻量级 FEC、物理层重传(PLR)和 UPHY 恢复实现无损传输,链路层采用基于信用的流控(CBFC)消除丢包,系统层消除单点故障,软件层提供 Dynamo Shadow Engine Recovery 和检查点恢复。NVIDIA 称该架构可实现比通用以太网低 3 倍的端到端延迟和 10 倍的数据包速率,并支撑 Vera Rubin NVL72 将 72 颗 Rubin GPU 连接为单一计算域。
正文节选
For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster must synchronize gradients across thousands of collective operations per second. Similarly, during inference, unplanned downtime directly reduces the total volume of requests served, strictly limiting revenue generation. As AI models grow exponentially, the network infrastructure required to train and serve them must scale in tandem. Howeve