ToolArtist:用于智能体图像生成的工具使用统一多模态模型
原标题:ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
AI 摘要
ToolArtist 提出了一种完全智能体的图像生成模型,通过对统一多模态模型进行后训练,将推理、外部工具调用和原生图像生成整合到单一策略中。在监督微调阶段,教师智能体使用搜索和图像生成工具,并将轨迹转换为统一多模态模型兼容格式;在强化学习阶段,引入 Reason-Act-Draw GRPO 算法,结合意图和质量奖励进行联合优化。实验表明,将整个开放世界图像生成过程置于智能体策略下,优于固定流程或部分智能体控制的方法。
正文节选
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Abstract Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescribe a fixed workflow or place only a subset of the open-world image generation process under ag