NVIDIA Groq 3 LPX 确定性执行如何驱动高能效推理
原标题:How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
AI 摘要
NVIDIA 在技术博客中介绍其 Vera Rubin 平台如何通过功耗管理实现高能效 AI 推理,核心是 NVIDIA Groq 3 LPX 的确定性执行模型。该模型让 LPU 编译器在运行前生成精确到时钟周期的执行调度,从而预测每个周期的电流需求,并借助 Preemptive Power(PEP)和 Clock Period Synthesis(CPS)两项技术降低电压保护带,使更多电力用于实际 AI 负载。文章还介绍了工厂级 DSX MaxLPS 软件和机架级电容与智能功耗平滑软件,称可在同一站点电力范围内多部署最多 40% 的 GPU 并提升 35% 的 token 吞吐。
正文节选
Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize output within the factory’s limited power budget. This makes performance per watt—rather than raw, unnormalized throughput—the ultimate measure of an AI platform’s value. The NVIDIA Vera Rubin platform is designed to enable power-efficient AI at scale. At its core is NVIDIA Vera Rubin NVL72, which delivers strong performance per watt across