视觉主导与延迟方法:提升VLM个性化安全
原标题:When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
AI 摘要
该研究提出MPS-Bench基准,包含5181个来自584张真实图像的高风险场景,用于评估视觉语言模型(VLM)的个性化安全。研究发现8个前沿VLM在86-99%的情况下直接回答而非寻求缺失上下文,个性化安全得分均不超过2.6/5。通过因果干预,研究者识别出'视觉主导'机制:视觉信息在早期层进入文本表示并抑制文本风险信号,导致后期内部修复不可靠。为此提出PRISM输入监控器,采用双向跨模态调制预测何时需要延迟回答,AUC达0.978,在所有测试模型上严格优于安全-效用帕累托前沿。
正文节选
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs Abstract Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reasonable in general may still be unsafe for a particular user whose medical, emotional, or situational context is unknown to the model. We study this problem of personalized safety in multimodal systems and introduce MPS-Bench, a benchmark of 5,181 scenarios from 584 real-worl