NVIDIA 在 GB300 NVL72 上支持 Qwen3.8-2.4T-A95B 模型部署
原标题:Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
AI 摘要
阿里巴巴发布了其最大的开源权重模型 Qwen3.8-2.4T-A95B(Qwen3.8-Max),总参数达 2.4T,每个 token 激活 95B 参数,采用细粒度 MoE 架构和全注意力与线性注意力混合设计,支持百万 token 上下文。NVIDIA 宣布该模型可在 GB300 NVL72 平台上以 FP8 精度实现每 GPU 每秒超 4K tokens 的吞吐量,并通过 SGLang、vLLM、NVIDIA Dynamo 和 NIM 等推理栈提供部署支持。该模型专为智能体工作负载设计,内置推理控制选项,并可通过 NeMo AutoModel 进行微调。
正文节选
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem. It has 2.4T total parameters with 95B activated per token. It has 2.4T total parameters with 95B activated per token. It’s a fine-grained mixture of experts (MoE) architecture with a hybrid of full and linear attention, a context window of up to one million tokens, and an output length of up to 128K, designed for demanding reasoning and