返回全部动态

AWS 教程:在 SageMaker HyperPod 上部署 Qwen3.8-2.4T-A95B

原标题:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AWS Machine Learning Blog一手来源教程质量 88

AI 摘要

AWS 机器学习博客发布教程,介绍如何在 Amazon SageMaker HyperPod 上使用 vLLM 部署阿里 Qwen 团队开源的 Qwen3.8-2.4T-A95B 模型。该模型总参数 2.4 万亿,激活参数 950 亿,采用混合线性注意力与全注意力架构,原生上下文 262K,支持 NVFP4 量化、推理控制和多 token 预测。文章详细说明了架构、能力、基准测试及 HyperPod 的部署优势,为自托管万亿参数模型提供参考。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM On August 12, 2026, Alibaba’s Qwen team released Qwen3.8-2.4T-A95B. This is the first time a Qwen-Max-class model has been made available as open weights. With 2.4 trillion total parameters (95 billion activated per token), a hybrid linear-plus-full-attention architecture, and native context up to 262K tokens (extensible to 1M), Qwen3.8 targets the most demanding agentic and reasoning workloads. These include multi-step coding, l


发布时间:2026-09-10 06:26
抓取时间:2026-09-10 06:28
来源机构:AWS
阅读原文aws.amazon.com