arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向计算溯源:在生成文本中携带因果状态证据

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

Benjamin Belay

arXiv 2608.16868首次发表:更新:

AI 中文总结

该研究提出计算溯源理念,通过在两种受控模型架构中让已验证内部状态决定生成文本的统计模式,证明答案不变时仍可在生成文本中保留因果相关内部状态的信息。

AI 中文摘要

语言模型的输出本身无法为产生该输出的内部计算提供可验证的证据。我们研究计算溯源:生成文本是否能携带可检测的、与因果相关的内部状态的证据。我们在两种受控架构中测试这一理念的受限形式:模块化前馈神经网络和基于Transformer的模型。两种架构均在同一算术任务上训练,且必须经过两个离散的中间状态,允许不同内部路径产生相同答案。我们在这些路径间刻意切换,验证实际使用的状态,并让该已验证状态决定生成文本中的一种细微统计模式,该模式可在后续被检测到。前馈系统和Transformer系统在公开及单独密封的受保护端到端评估中,均通过了全部128个匹配对,检测器成功恢复了与已验证内部状态相关的信号。所需的因果计算在5个独立训练的前馈模型和3个独立训练的Transformer中均复现。在仅输出答案的Transformer实验中,我们的线性探针未恢复自然学习到的中间状态。这些结果提供了受控的概念验证:即使答案不变,已验证的、与因果相关的内部状态的信息也可保留在生成文本中。

英文摘要

A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a modular feed-forward neural network and a transformer-based model. Both architectures are trained on the same arithmetic task with a mandatory pathway through two discrete intermediate states, allowing different internal paths to produce the same answer. We deliberately switch between these paths, authenticate the state actually used, and let that verified state determine a subtle statistical pattern in the generated text that can later be detected. The feed-forward and transformer systems each passed all 128 matched pairs in both their public and separately sealed protected end-to-end evaluations, with the detector recovering the signal associated with the authenticated internal state. The required causal computation also reproduced across five independently trained feed-forward models and three independently trained transformers. In a separate answer-only transformer experiment, our linear probes did not recover a naturally learned intermediate state. These results provide a controlled proof of concept that information about a verified, causally relevant internal state can be preserved in generated text even when the answer is unchanged.

Comments16 pages, 1 figure, 7 tables

Journal refNeurIPS 2026 Workshop on Interpretability as a Science

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑