arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

定向幻觉:基于新闻的语言模型问答中的意识形态漂移

Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering

Chendi Wang, Liam Cunningham, Tom Yishay, Jieying Chen

arXiv 2607.20487首次发表:更新:

发表机构

Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于新闻的语言模型问答中的意识形态漂移,提出可重复测量框架,利用大量新闻文章对多个模型进行实验,发现幻觉率因模型和话题而异,幻觉内容有向左漂移现象,并探讨了相关意义。

AI 中文摘要

大语言模型(LLMs)越来越多地用于回答政治信息相关问题,在选举相关信息场景中,事实错误和意识形态扭曲风险很高。我们提出了一个可重复的测量框架,将基于文档的问答中出现的幻觉(无支撑陈述)视为意识形态漂移的诊断信号。利用来自QBias的21727篇由专家标注的美国政治新闻文章,涵盖左、中、右不同来源,我们进行了一系列操作,包括生成特定文章问题、从多个模型获取基于文档的答案、检测句子级幻觉、用微调的立场分类器对幻觉句子的意识形态价进行分类以及探究输出对数来关联令牌级不确定性与幻觉和漂移。幻觉率因模型而异且集中在有争议话题,幻觉内容呈现出明显的向左漂移。我们还讨论了对审核人工智能介导的政治信息以及在选举相关部署中设计保障措施的意义。

英文摘要

Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent information settings where factual errors and ideological distortions are high-stakes. We present a reproducible measurement framework that treats hallucinations, unsupported statements in document-grounded QA, as diagnostic signals of ideological drift. Using 21,727 expert-labeled U.S. political news articles from QBias spanning left, center, and right sources, we (i) generate an article-specific question, (ii) elicit document-grounded answers from three open-weight LLMs and one proprietary model, (iii) detect sentence-level hallucinations via reference-based comparison, (iv) classify the ideological valence of hallucinated sentences with a fine-tuned stance classifier, and (v) probe output logits to relate token-level uncertainty to hallucination and drift. Hallucination rates vary substantially across models and concentrate in contentious topics, while source-ideology differences in hallucination frequency are modest. In contrast, hallucination content exhibits robust leftward drift: a majority of hallucinated sentences are classified as left-leaning, including among hallucinations generated from right-leaning sources. Logit-level analysis shows hallucinations arise in high-entropy generation contexts, and in some models uncertainty also predicts leftward drift, consistent with an "uncertainty to guessing" mechanism. We discuss implications for auditing AI-mediated political information and for designing safeguards in election-relevant deployments.

DOI:10.24963/ijcai.2026/831

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑