返回全部动态

通用多模态基础模型:基于合成数据与上下文学习

原标题:Generalized Multimodal Foundation Model

arXiv cs.LG一手来源研究质量 80

AI 摘要

该论文提出一种通用多模态基础模型,旨在解决现有融合模型只能处理预定义模态和单一任务的问题。作者基于结构多模态因果模型(SMCM)构建大规模合成多模态数据集进行训练,使模型学习可迁移的多模态关联模式,而非依赖特定模态。模型扩展自表格基础模型 TabPFN,通过统一 token 生成机制将异构模态映射到统一空间,并在推理时利用上下文示例激活任务相关关联,无需参数更新。在多个真实数据集上,该模型无需任务特定适配即可达到与专用模型相当的性能。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Generalized Multimodal Foundation Model Abstract Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single tasks, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet rather aggressive question arises, whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations


发布时间:2026-09-22 12:00
抓取时间:2026-09-22 12:05
来源机构:arXiv
阅读原文arxiv.org