无辜画面,仇恨故事:多轮视觉故事生成中的仇恨意图评估与检测
原标题:Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
AI 摘要
该研究指出,前沿文生图系统(如Gemini和GPT-Image)支持多轮对话生成一致的角色和场景,使得制作仇恨性视觉故事(有序图像组)变得廉价且可扩展。作者引入了HatefulStoryPrompts数据集,包含55个仇恨故事、330个多轮配置,并评估了五个前沿模型,发现所有模型都能完成超过80%的故事,最强模型达到99.0%。现有审核系统在检测组级仇恨含义时表现不佳,专用安全模型召回率最高仅34.9%,强视觉语言模型为67.5%。作者提出了交互感知监控和生成后防御方法,分别达到97.3%和92.6%的召回率,以及80.2%的生成后检测召回率。
正文节选
Computer Science > Computer Vision and Pattern Recognition Title:Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation View PDF HTML (experimental) Abstract:Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emph{Der Giftpilz}. Recently, frontier text-to-image (T2I) systems such as Gemi