分布式隐性危害:多模态大模型视频审核的组合安全盲点
原标题:Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
AI 摘要
该研究提出“分布式隐性危害”(DIH)概念,指多模态大模型在视频审核中无法识别由看似无害的片段组合而成的有害内容。作者构建了包含9000多个视频的DIH数据集和基准,测试30多个MLLM后发现即使最强模型检测率也低于45%,真实视频中同样失败。通过两阶段后训练,检测准确率提升超60个百分点,表明该缺陷可通过针对性训练缓解。
正文节选
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation Abstract Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a compositional safety blind spot: videos composed of seemingly benign components can convey harmful meaning when interpreted as a whole. We refer to this phenomenon as Distributed Implicit Harm (DIH), where harm arises from relations among components distributed along a decomposition axis of the video