发表机构
Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过68个LLM生成的侦探悬疑对话和369人实验,发现锚定、谬误忽视等修辞策略能显著操纵人类裁决,使AI安全辩论的监督机制失效。
AI 中文摘要
基于人类反馈的强化学习(RLHF)在使大型语言模型响应人类指令方面发挥了核心作用。然而,人类评估者往往偏爱奉承或具有说服力的回应,而非真实准确的回应,这促使模型以牺牲准确性为代价来迎合评估者。AI安全辩论被提出作为一种改进语言模型监督的方式:在这种范式中,两个智能体就对立立场进行辩论并质疑对方的论断,从而可能向裁决者揭示虚假信息。AI安全辩论的一个核心前提是,在对抗性审查下,真实的论点比虚假的论点更容易辩护。在本工作中,我们研究了当辩手利用人类判断中的偏见采用修辞策略时,这一优势是否仍然存在。受竞技辩论启发,我们构建了68个由LLM生成的关于已知罪犯的侦探悬疑故事的对话,涵盖四种干预措施:锚定效应、谬误忽视、支持行话和冗长表述。我们将每种干预分别应用于支持真实罪犯的一方或支持无辜嫌疑人的一方,从而能够区分对裁决的影响与正确性的影响。在一项涉及369名参与者的研究中,我们发现,汇总各类偏见后,这些干预显著地将判断转向被操纵的一方。这些发现暴露了基于辩论的监督的一个脆弱性:人类裁决对操纵性修辞策略敏感。
英文摘要
Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue opposing positions and challenge each other's claims, potentially exposing falsehoods to the adjudicator. A central premise of AI safety via debate is that truthful arguments are easier to defend than false ones under adversarial scrutiny. In this work, we investigate whether this advantage persists when debaters use rhetorical strategies that exploit biases in human judgment. Inspired by competitive debate, we construct 68 LLM-generated dialogues about detective mysteries with known culprits, spanning four interventions: anchoring, fallacy oversight, pro-jargon, and verbosity. We apply each intervention to either the side advocating for the true culprit or the side advocating for an innocent suspect, allowing us to distinguish influence on adjudication from correctness. In a study with 369 participants, we find that, pooled across bias types, these interventions significantly shift judgments toward the manipulated side. These findings expose a vulnerability in debate-based supervision: human adjudication is sensitive to manipulative rhetorical strategies.