arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05114cs.CL

信念轨迹能量:衡量通向预测的路径

Belief-Trajectory Energy: Measuring the Path to a Prediction

发表机构复旦大学 · 中国科学技术大学 · 上海创新研究院
查看机构详情
  • Fudan University(复旦大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

Jiahao Ying, Wei Tang, Boxian Ai, Yaoning Wang, Haotian Chen, Wenhe Sun, Caijun Xu, Haozhan Cai, Changyi Xiao, Yixin Cao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出信念轨迹能量(BTE),通过度量LLM各层预测修正来刻画输入,在推理难度、人机审查检测和生成器归因上表现优异,并揭示其反映模型学习内容。

中文摘要 AI 辅助

大型语言模型(LLMs)在Transformer各层中逐步修正其预测,然而我们通常只观察最终输出,忽略了其形成过程中的轨迹。我们引入了信念轨迹能量(BTE),这是一种基于模型的度量,通过模型中层级的预测修正来刻画输入。通过将中间状态映射到共享的预测空间,BTE提供了一种原则性的信念变化度量,可概括为标量或结构化深度剖面。理论上,我们证明了局部BTE对应于Fisher-Rao几何下的预测修正,而修正序列捕获了超越初始到最终信念变化的信息。实证上,标量BTE在各种推理任务中提供了模型相对的难度信号,而更丰富的BTE表示支持人机LLM审查检测和细粒度生成器归因,达到了高达0.998的宏AUROC和95.6%的八路归因准确率。进一步分析表明,BTE在预训练过程中发展,并通过针对性训练被选择性重塑,表明所得到的度量反映了评分模型所学到的内容。总之,我们的结果确立了信念轨迹作为原则性的模型基础信号,并提出了一种更广泛的视角,即学习模型本身可以作为表征其所处理数据的工具。更多演示可在该https URL找到。

英文摘要

Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to $0.998$ macro-AUROC and $95.6\%$ eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at https://yingjiahao14.github.io/BTE-web/.

↑