返回全部动态

融合单模态与视觉语言表征的胸部X光多标签分类

原标题:Integrating Unimodal and Vision-Language Representations in Latent Space for Multi-Label Chest X-Ray Classification

arXiv cs.CV一手来源研究质量 75

AI 摘要

该研究提出一个融合框架,将RAD-DINO的单模态视觉表征与BioViL-T的视觉-语言表征在潜在空间中分别精炼后归一化并跨三分支融合,用于MIMIC-CXR-JPG数据集上14类胸部X光多标签分类。实验显示RAD-DINO单独使用时优于BioViL-T,而早期融合进一步提升结果,表明两种嵌入来源包含互补信息;混合融合在潜在空间精炼后相比早期融合有一致且统计显著的提升。研究仅在MIMIC-CXR-JPG上内部评估,跨机构泛化性仍待验证,源代码已公开。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Integrating Unimodal and Vision–Language Representations in Latent Space for Multi-Label Chest X-Ray Classification Abstract Multi-label chest X-ray classification has attracted considerable attention in recent years, with the effective use of visual representations and clinical semantic knowledge playing an important role. This study proposes a framework that combines unimodal representations from RAD-DINO with vision–language representations from BioViL-T for the classification of 14 labels in


发布时间:2026-09-10 12:00
抓取时间:2026-09-10 12:14
来源机构:arXiv
阅读原文arxiv.org