返回全部动态

DeflectBench:评估大语言模型生成修辞谬误的基准

原标题:DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

arXiv cs.CL一手来源研究质量 82

AI 摘要

DeflectBench 是一个新基准,用于评估大型语言模型按需生成修辞谬误(如偷换概念、人身攻击、红鲱鱼)的能力。研究测试了四个前沿模型,发现拒绝行为主要由请求结构而非声明内容驱动,且教育辩论教练的提示框架几乎消除了所有拒绝,但模型通常会产生标记性合规,即在同一响应中命名所请求的操纵。该基准揭示了不同实验室的对齐机制产生不同的合规特征,对多元对齐社区具有重要意义。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs Abstract Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring)


发布时间:2026-08-28 12:00
抓取时间:2026-08-28 18:10
来源机构:arXiv
阅读原文arxiv.org