arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28673cs.CL

基准测试大语言模型的论证行为:针对人身攻击的防御策略研究

Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks

  • Warsaw University of Technology(华沙理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Ewelina Gajewska, Katarzyna Budzynska, Jaroslaw Chudziak

AI总结:

本研究通过对比美国总统辩论语料库,发现大语言模型在应对人身攻击时僵化采用逻辑防御,缺乏人类辩手的人格反击策略,归因于安全微调限制了其策略空间。

AI中文摘要:

大语言模型(LLMs)越来越多地被部署为说服性对话中的论证智能体,因此有必要对其相对于人类对话者的辩论能力进行严格评估。在本研究中,我们聚焦于人身攻击(ad hominem 论证),这类论证传统上被视为谬误,但在政治说服性对话中却扮演着关键角色,其中人格(ethos)往往与命题内容并驾齐驱。具体而言,我们探究现代大语言模型能否复现人类在策略性使用和回应此类攻击方面的能力。我们分析了一个自然语言政治对话语料库,以识别人类对话者在以人格为中心的辩论中自然采用的防御策略,并将其构建为一个对话游戏。在实证层面,我们将大语言模型生成的对话与美国总统辩论的 Elec60to16-fallacy 语料库进行基准对比,将人类辩手的防御策略库与人工智能体的策略库进行对照。结果显示存在显著差异:大多数大语言模型僵化地优先采用逻辑防御,未能将人格反击作为政治话语中的有效策略加以利用。我们认为,当前的安全微调限制了大语言模型的策略行动空间,使其无法在人格争议被视为规范性预期而非单纯谬误的领域中充分参与自然交互。

英文摘要:

Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.

补充信息

↑