返回全部动态

DeepSeek-V3 发布:高效 MoE 模型,性能比肩闭源

原标题:deepseek-ai/DeepSeek-V3 README.md

DeepSeek Harness Repository Activity一手来源模型发布质量 88

AI 摘要

DeepSeek 发布了 DeepSeek-V3,这是一个拥有 6710 亿总参数、每个 token 激活 370 亿参数的混合专家(MoE)语言模型。该模型采用无辅助损失负载均衡策略和多 token 预测训练目标,在 14.8 万亿 token 上预训练,仅需 278.8 万 H800 GPU 小时,性能超越其他开源模型,并可与领先闭源模型媲美。训练过程稳定,无不可恢复的损失尖峰。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

<!-- markdownlint-disable first-line-h1 --> <!-- markdownlint-disable html --> <!-- markdownlint-disable no-duplicate-header --> <div align="center"> <img src="https://github.com/deepseek-ai/DeepSeek-V2/blob/main/figures/logo.svg?raw=true" width="60%" alt="DeepSeek-V3" /> </div> <hr> <div align="center" style="line-height: 1;"> <a href="https://www.deepseek.com/"><img alt="Homepage" src="https://github.com/deepseek-ai/DeepSeek-V2/blob/main/figures/badge.svg?raw=true"/></a> <a href="ht


发布时间:2025-06-16 14:34
抓取时间:2026-08-02 00:09
来源机构:DeepSeek
阅读原文github.com