返回全部动态

Ovis-Embedding:推进通用全模态嵌入前沿

原标题:Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

arXiv cs.AI一手来源研究质量 90

AI 摘要

该论文提出 Ovis-Embedding,一个基于 Qwen-Omni 预训练理解模型构建的全模态嵌入模型家族,通过共享多模态骨干网络将文本、图像、视频和音频编码到统一表示空间。研究采用低秩对比预训练、同源采样、focal loss 和嵌入蒸馏等训练优化策略,并在推理时用低秩特征分解实现紧凑嵌入。在 MMEB-v3、MMEB-v2、MVEB、MAEB 和 RTEB 等基准上取得最优性能,展示了统一全模态训练在任意模态检索任务中的潜力。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Abstract In this report, we introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make three key advances: (1) native omni-modal initialization: we adopt a pretrained


发布时间:2026-09-24 12:00
抓取时间:2026-09-23 12:12
来源机构:arXiv
阅读原文arxiv.org