返回全部动态

Moonworks Lunara:建模艺术智能的文本到图像模型

原标题:Moonworks Lunara: Modeling Artistic Intelligence

arXiv cs.CV一手来源模型发布质量 72

AI 摘要

Moonworks 提出 Moonworks Lunara 文本到图像模型,将「艺术智能」定义为先建立语义、艺术与构图约束、再在约束内实现视觉世界的两阶段计算,并采用新型 Diffusion Mixture Transformer 架构与 CAT 训练算法。团队用 GPT-5.6 Sol 作为评估器,在 1000 条提示、8000 张生成图像上对比包括 FLUX.2-Klein-4B、Qwen-Image、GPT-Image-1-Mini 在内的七个模型,Lunara 在美学质量上排名第一、情感共鸣第二,盲测人类评估也排名第一。该模型活跃参数低于 10B,1024x1024 图像在 40GB A100 上推理延迟低于 10 秒。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Moonworks Lunara: Modeling Artistic Intelligence Abstract We formulate Artistic Intelligence as exploration driven world realization, leaving space for creative possibility while preserving the semantic, artistic, and compositional structure that must remain true. Moonworks Lunara, a text-to-image model, implements this framework with a novel Diffusion Mixture Transformer architecture. A new training algorithm iteratively evolves the data distribution through informative sample acquisition and t


发布时间:2026-09-23 12:00
抓取时间:2026-09-22 12:32
来源机构:arXiv
阅读原文arxiv.org