NVIDIA Exemplar Cloud:解锁 AI 基础设施全性能的配置诊断经验
原标题:NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure
AI 摘要
NVIDIA 技术博客指出,相同硬件配置的 AI 集群因内核、虚拟机、BIOS 和 NCCL 等配置差异,训练吞吐量可相差 8%-12%。文章通过四个案例展示了如何利用 perf、Nsight Systems 等工具定位并解决性能瓶颈,例如启用 CMDQV/VCMDQ 修复虚拟化环境中的 SMMU 开销问题。这些诊断模式可帮助基础设施工程师在正式验证前优化集群性能。
正文节选
Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We routinely see 8% to 12% gaps between partner deployments and the corresponding NVIDIA reference architecture (RA) on the same workload, same model, same global batch size. The cause is often a stack of configuration choices in the kernel, hypervisor, BIOS, and NVIDIA Collective Communications Library (NCCL) settings, each costing a few percent,