返回全部动态

LLaDA-Image:全开放训练配方的强图像生成模型

原标题:Paper page - LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

LLaDA-Image 是一个将 6B 扩散 Transformer 与冻结的视觉语言模块相结合的统一框架,通过仅图像预训练和 Muon 优化器生成逼真图像并支持精确编辑。该模型被蒸馏为 LLaDA-Image-Turbo,可在 2-4 步内快速推理,在 Qwen-Image-Bench 的英文和中文赛道均创下开源模型最佳成绩。研究团队已发布模型权重、训练代码和详细配方,以支持进一步研究。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Abstract LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results. We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a froz


发布时间:—
抓取时间:2026-09-04 10:42
来源机构:Hugging Face
阅读原文huggingface.co