微信发布多模态嵌入模型WeMM-Embedding,刷新多项基准
原标题:Paper page - WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
AI 摘要
腾讯微信视觉团队发布了WeMM-Embedding,一个支持文本、图像、视频和交错输入的多模态嵌入模型系列,包含2B、4B和9B三种规模。该模型在MMEB-v2基准上,2B版本已超越此前领先的8B开源基线,9B版本以80.6分创下新纪录。模型已在微信视频号、公众号、朋友圈和电商等场景大规模部署,并在14个在线A/B测试中表现一致提升。模型权重和代码已在GitHub开源。
正文节选
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Abstract WeMM-Embedding is a family of universal multimodal embedding models that align text, images, videos, and interleaved inputs in a shared space, achieving state-of-the-art retrieval and recommendation performance across public benchmarks and large-scale WeChat applications. Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for a