返回全部动态

NVIDIA 用 QAD 开发 Nemotron 3.5 Lightning NVFP4 模型

原标题:Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer

NVIDIA Technical Blog一手来源研究质量 85

AI 摘要

NVIDIA 发布了 Nemotron 3.5 Lightning NVFP4 检查点,通过量化感知蒸馏(QAD)将模型从 66 GB 压缩至 22 GB,同时实现高达 4 倍的吞吐量提升。该过程使用 NVIDIA Model Optimizer,先进行 PTQ 量化,再通过蒸馏恢复精度。文章详细介绍了 QAD 的两阶段流程、PTQ 配方选择及评估结果,表明 QAD 能有效恢复激进量化带来的精度损失。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models, developers can find the right-sized model for their needs. The new Nemotron 3.5 Lightning NVFP4 checkpoint, for example, preserves accuracy while unlocking up to 4x faster throughput. It’s compressed down to 22 GB from the 66 GB full precision checkpoint by quantizing many of its weights to 4 bits. To compress models to NVFP4, post-training quantization (PTQ)


发布时间:2026-08-18 02:12
抓取时间:2026-08-18 02:46
来源机构:NVIDIA
阅读原文developer.nvidia.com