FLUX 3 发布:统一多模态流模型,支持图像视频音频与动作预测
原标题:FLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action Prediction
AI 摘要
Black Forest Labs 发布了 FLUX 3,这是一个统一的多模态流匹配基础模型,支持图像、视频、音频和动作预测。其核心创新是 Self-Flow 框架,通过自监督特征重建目标同时提升生成质量和表征质量,在机器人操作任务中成功率从 42% 提升至 71%。FLUX 3 支持长达 20 秒的同步音频视频生成,并计划后续开放 Dev 权重。
正文节选
| Field | Value | |---|---| | Developer | Black Forest Labs | | Model type | Multimodal flow matching foundation model (Diffusion Transformer) | | Modalities | Image, Video, Audio, Action Prediction (text-conditioned) | | Training framework | Self-Flow (self-supervised flow matching) | | Training data | Tens of millions of hours of general video, hundreds of thousands of hours of manipulation video, plus images and audio | | Video length | Up to 20 seconds with synchronized audio; early ev