FailSAE:利用稀疏自编码器实现视觉语言模型的可解释失败预测
原标题:FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders
AI 摘要
Wayne State University的研究人员提出FailSAE框架,利用稀疏自编码器(SAE)对视觉语言模型(如CLIP)进行可解释的失败预测。该框架将失败预测建模为对SAE稀疏潜在激活的分类任务,并引入三阶段失败感知训练流程。实验表明,FailSAE在失败预测上优于现有基线,并揭示了失败时模型表示从类别特定概念转向模糊或风格相关概念。此外,该框架还支持运行时失败恢复。
正文节选
FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders Abstract Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by aligning visual and textual representations in a shared embedding space. As VLMs are increasingly used for high-stakes domains, failure prediction becomes critical for risk-aware deployment and human intervention. Existing failure prediction methods typically rely on confidence scores