LLaDA-Image:全开放训练配方的强图像生成模型
原标题:Paper page - LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
AI 摘要
LLaDA-Image 是一个将 6B 扩散 Transformer 与冻结的视觉语言模块相结合的统一框架,通过仅图像预训练和 Muon 优化器生成逼真图像并支持精确编辑。该模型被蒸馏为 LLaDA-Image-Turbo,可在 2-4 步内快速推理,在 Qwen-Image-Bench 的英文和中文赛道均创下开源模型最佳成绩。研究团队已发布模型权重、训练代码和详细配方,以支持进一步研究。
正文节选
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Abstract LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results. We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a froz