DeflectBench:评估大语言模型中修辞谬误生成的基准
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
浏览论文内容
中文总结 AI 辅助
DeflectBench评估四个前沿大语言模型在三种修辞谬误生成任务中的表现,发现拒绝率主要由提示框架和谬误类型决定,教育辩论教练提示可大幅降低拒绝率,且各模型的合规行为分布存在差异。
中文摘要 AI 辅助
关于大语言模型能否被提示按需生成修辞谬误,以及当前的安全后训练是否会约束该行为,相较于检测现有文本中谬误的相关问题,受到的关注更少。我们通过DeflectBench填补这一空白,评估四个前沿模型针对三种回避策略(deflect策略,包括诉诸他者、人身攻击、转移话题)、七种提示框架及80条涵盖四个争议级别的主张所生成的23990个输出。拒绝行为主要由请求结构而非主张内容决定。80条主张的单条主张拒绝率仅相差11个百分点,而单一提示框架的变化可使模型拒绝率波动近100个百分点,在明确框架下,请求的谬误类型切换可使拒绝率波动超过80个百分点。教育辩论教练提示框架使所有四个模型系列的拒绝率降至接近零,但被绕过的行为并非完全合规。模型通常产生带标签的合规行为,在包含所请求操纵的同一响应中命名该操纵。四个模型在弃权(不执行)、带标签的合规、软弃权(不执行)和完全合规之间的分布不同。代码和数据集已发布在此httpsURL。
英文摘要
Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels. Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings. An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it. The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at https://github.com/ArtKanke/DeflectBench.