返回全部动态

FLUX 3 发布:统一多模态流模型,支持图像视频音频与动作预测

原标题:FLUX 3 Model Overview: Multimodal Flow Models for Image, Video, Audio, and Action Prediction

Hugging Face Blog一手来源模型发布质量 88

AI 摘要

Black Forest Labs 发布了 FLUX 3,这是一个统一的多模态流匹配基础模型,支持图像、视频、音频和动作预测。其核心创新是 Self-Flow 框架,通过自监督特征重建目标同时提升生成质量和表征质量,在机器人操作任务中成功率从 42% 提升至 71%。FLUX 3 支持长达 20 秒的同步音频视频生成,并计划后续开放 Dev 权重。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

| Field | Value | |---|---| | Developer | Black Forest Labs | | Model type | Multimodal flow matching foundation model (Diffusion Transformer) | | Modalities | Image, Video, Audio, Action Prediction (text-conditioned) | | Training framework | Self-Flow (self-supervised flow matching) | | Training data | Tens of millions of hours of general video, hundreds of thousands of hours of manipulation video, plus images and audio | | Video length | Up to 20 seconds with synchronized audio; early ev


发布时间:
抓取时间:2026-08-03 01:40
来源机构:Hugging Face
阅读原文huggingface.co