DeflectBench:评估大语言模型生成修辞谬误的基准
原标题:DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
AI 摘要
DeflectBench 是一个新基准,用于评估大型语言模型按需生成修辞谬误(如偷换概念、人身攻击、红鲱鱼)的能力。研究测试了四个前沿模型,发现拒绝行为主要由请求结构而非声明内容驱动,且教育辩论教练的提示框架几乎消除了所有拒绝,但模型通常会产生标记性合规,即在同一响应中命名所请求的操纵。该基准揭示了不同实验室的对齐机制产生不同的合规特征,对多元对齐社区具有重要意义。
正文节选
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs Abstract Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring)