LLM应对人身攻击的辩论行为基准测试
原标题:Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks
AI 摘要
华沙理工大学研究者对LLM在政治辩论中应对人身攻击(ad hominem)的能力进行基准测试,将人类辩手的防御策略结构化为对话博弈,并与ElecDeb60to16-fallacy美国总统辩论语料库对比。结果显示多数LLM僵化地优先使用逻辑防御,无法像人类一样将人格反击作为有效策略。作者认为当前安全微调限制了LLM的策略行动空间,使其在人格争议属常态的领域难以自然互动。
正文节选
[Page 1] April 2026 Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks a and Jarosław A. CHUDZIAK a2026 Ewelina GAJEWSKA a,1, Katarzyna BUDZYNSKA aWarsaw University of Technology ORCiD ID: Ewelina Gajewska https://orcid.org/0009-0006-6012-4