ProtoLIP:从句子级到对象级证据解耦
原标题:ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
AI 摘要
论文提出 ProtoLIP,一种轻量级原型中介证据层,用于解决视觉语言模型中对象级查询仍保留共现对象和共享上下文证据的问题。ProtoLIP 将可复用视觉原型组织为文本衍生的语义族,并通过查询依赖的族路由约束哪些原型可提供证据,无需空间标注也不重训骨干网络。在冻结的 ItemizedCLIP 上仅增加 1.48% 可训练参数,Pointing 和 Energy 分别最高提升 42.2% 和 55.1%,且定位增益可迁移到独立预训练的 VLM。
正文节选
ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement Abstract Query-conditioned vision–language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model’s prediction. Across multiple VLM architectures and independent benchmar