返回全部动态

Skim and Skip:高效多模态检索的分层自适应推理框架

原标题:Skim and Skip: Hierarchical Adaptive Inference for Efficient Multimodal Retrieval

arXiv cs.IR一手来源研究质量 84

AI 摘要

清华大学、微软和上海交通大学的研究者提出了 Skim and Skip (SAS) 框架,用于提升多模态检索的效率。该框架通过基于 [EOS] 的 token 级信息选择和自适应深度推理,在 12 个 MMEB 检索任务上保留了约 99% 的密集基线性能,同时实现了最高 1.64 倍的端到端加速和 66.3% 的 FLOPs 减少。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Skim and Skip: Hierarchical Adaptive Inference for Efficient Multimodal Retrieval Abstract Universal multimodal retrieval (UMR) increasingly adopts multimodal large language models (MLLMs) as unified embedding backbones, but their strong retrieval performance comes at substantial inference cost. Existing methods typically rely on uniformly dense inference, where all input tokens are processed through the entire model and matched using the final-layer [EOS] representation. However, this paradigm


发布时间:2026-09-03 12:00
抓取时间:2026-09-03 12:56
来源机构:arXiv
阅读原文arxiv.org