Mistral AI 发布 Pixtral 12B 多模态模型
原标题:ResearchAnnouncing Pixtral 12BSeptember 17, 2024By Mistral AI team
AI 摘要
Mistral AI 发布了 Pixtral 12B,这是一个原生多模态模型,基于 Mistral Nemo 架构,并新训练了 4 亿参数的视觉编码器,支持可变图像尺寸和 128K 长上下文。该模型在 MMMU 基准上达到 52.5%,在指令跟随和多模态任务上表现优异,同时保持文本性能不妥协。模型采用 Apache 2.0 许可证,可在 La Plateforme 和 Le Chat 上使用。
正文节选
Pixtral 12B in short: - Natively multimodal, trained with interleaved image and text data - Strong performance on multimodal tasks, excels in instruction following - Maintains state-of-the-art performance on text-only benchmarks - Architecture: - New 400M parameter vision encoder trained from scratch - 12B parameter multimodal decoder based on Mistral Nemo - Supports variable image sizes and aspect ratios - Supports multiple images in the long context window of 128k tokens - Use: - Lic