线性判别树集成的可解释多模态分类
原标题:Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles
AI 摘要
该研究提出一种基于线性判别树集成(LDT、LDF、LDAB)的可解释多模态分类框架,通过编码各模态为令牌、提取概念聚类、路由融合并通过改进的特征重要性指标解释趋势,在IEMOCAP、CMU-MOSI和自定义数学数据集上,相比多模态Transformer和可解释多模态路由(IMR),在F1-mod上提升4.3%、准确率提升3.0%,且特征重要性的人类标注一致性显著更高。
正文节选
spacing=nonfrench Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles Abstract: Multimodal affect and behaviour classifiers that fuse heterogeneous text, audio, and visual streams must simultaneously achieve competitive accuracy and produce human-understandable explanations of the cues driving their decisions—a dual objective that current high-capacity models, notably Transformers, only partially address. While Transformers attain strong predictive performance, their