AI 中文总结
针对大语言模型的对抗性政治偏见漏洞,提出思维链提示与直接偏好优化等策略,其中递归自校正方法可显著提升模型的政治中立性表现。
AI 中文摘要
随着大语言模型(LLMs)成为信息检索与摘要任务的核心支撑,确保其始终保持无党派性且不受政治偏见影响,是构建更安全、更可信人工智能(AI)的关键一步。当前的模型对齐范式,如基于人类反馈的强化学习(RLHF),可使LLMs遵循总体安全指令,但这类指令调优可能被对抗性提示注入利用,进而生成不安全内容;尤其现代对齐技术未将政治偏见作为有害且有偏见的内容专门针对。为解决LLMs的这一漏洞,我们提出使用思维链(CoT)提示与直接偏好优化(DPO)的缓解策略:利用公开的立法视频数据集,通过LLMs生成摘要,经对抗性提示注入偏见后,在专为政治摘要设计的四维尺度上评估性能。本文提出多种方法保护LLMs免受政治偏见注入;结果显示,所提递归自校正方法使所有模型的政治中立性李克特量表基线从2.14提升至4.56,证明其可在推理时有效缓解LLMs生成摘要中的政治偏见。
英文摘要
As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI). Current model alignment paradigms, such as reinforcement learning from human feedback (RLHF), make LLMs follow overarching safety instructions. However, this instruction tuning can be exploited via adversarial prompt injection and be used to generate unsafe content. In particular, political bias has not been specifically targeted by modern alignment techniques as harmful and biased content. To address this vulnerability of LLMs, we propose mitigation strategies using Chain of Thought (CoT) prompting and Direct Preference Optimization (DPO). Using a public dataset of legislative videos, we generate summaries using LLMs, inject bias via adversarial prompting and evaluate their performance on a four axis scale designed for political summarization. In this paper, we present different methods to shield LLMs against the injection of political bias. Our results demonstrate that the proposed Recursive Self-Correction approach raises model performance from a Political Neutrality Likert scale baseline of 2.14 to 4.56, averaged across all models, demonstrating effective inference-time mitigation of political bias in LLM-generated summaries.