arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在事实核查:通过激活工程检测虚假信息

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

Pedro T. Barcelos, Otávio Parraga, Marcelo M. Delucis, Lucas M. Fraga, Lucas S. Kupssinskü, Rodrigo C. Barros

arXiv 2608.06417首次发表:更新:

发表机构

PUCRS; Kunumi Institute(巴西天主教大学(里约格兰德 do 苏里); 库纳米研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于激活工程的虚假信息检测框架,通过对比激活引出潜在空间的虚假信息方向,在多类模型和基准上实现了优于部分基线的检测性能,为可解释性驱动的虚假信息检测提供了新方向。

AI 中文摘要

网络虚假信息的泛滥催生了对可扩展检测系统的需求。现有多数方法依赖表层语言特征或外部知识检索,而我们将真实性视为语言模型表示空间的几何属性。我们提出一种基于激活工程的虚假信息检测框架,该框架利用Transformer模型的潜在几何结构。我们遵循对比激活加法(Contrastive Activation Addition,CAA)的均值差原理,通过对比成对真实与虚假陈述的激活,在残差流中引出虚假信息方向。推理时,将未见主张的最后一个token激活投影到该方向,再将投影后的表示输入多层感知机(Multilayer Perceptron,MLP)进行分类。该过程无需微调主干模型、无需外部证据检索,也无需除用于估计方向的对比对之外的特定任务监督。我们在来自Gemma、Llama和Qwen系列的11个模型(参数规模从2.7亿到120亿)上,在三个事实核查基准(AVeriTeC、LIAR和FACTors)上评估该方法。虚假信息方向在不同模型规模和架构系列中均可恢复,且最后一个token投影在LIAR和FACTors上的表现与零样本和少样本提示基线相当或更优,较小模型的增益最大。在AVeriTeC上的性能更有限,我们将其归因于其基于证据的标注方案。这些发现证明,真实性在预训练语言模型的潜在空间中是一个结构化、线性可分的概念,并表明可解释性驱动的虚假信息检测可作为基于检索的管道的实用补充。代码可在该https URL获取。

英文摘要

The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property of a language model's representation space. We introduce a misinformation detection framework grounded in activation engineering, which leverages the latent geometry of transformer models. Our approach elicits a misinformation direction in the residual stream by contrasting activations from paired truthful and false statements, following the difference-in-means principle of Contrastive Activation Addition (CAA). At inference time, the last-token activation of an unseen claim is projected onto this direction, and the projected representation is fed to an Multilayer Perceptron (MLP) for classification. The procedure requires no fine-tuning of the backbone model, no external evidence retrieval, and no task-specific supervision beyond the contrastive pairs used to estimate the direction. We evaluate the method across 11 models from the Gemma, Llama, and Qwen families, ranging from 270M to 12B parameters, on three fact-checking benchmarks: AVeriTeC, LIAR, and FACTors. The falsehood direction is recoverable across model scales and architectural families, and last-token projection matches or surpasses zero-shot and few-shot prompting baselines on LIAR and FACTors, with the largest gains observed for smaller models. Performance on AVeriTeC is more limited, which we attribute to its evidence-grounded labeling scheme. These findings provide evidence that truthfulness is a structured, linearly separable concept in the latent space of pretrained language models, and point toward interpretability-driven misinformation detection as a practical complement to retrieval-based pipelines. The code is available on https://github.com/Malta-Lab/LaFaCt.

Comments13 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑