arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小型Transformer通过上下文不变机制追踪潜在共同原因的贝叶斯证据

Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism

Amir Mohammadpour, Michael Franke

arXiv 2609.35161首次发表:更新:

发表机构

University of Tübingen(蒂宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究小型Transformer如何通过上下文不变机制,在跨上下文泛化中实现潜在共同原因的贝叶斯证据积累,并分离模型机制与真实因果结构。

AI 中文摘要

我们对一种关于共同原因的贝叶斯推理形式如何在小型、可处理的Transformer中作为跨上下文泛化而出现进行了深入研究。在近期工作的基础上,我们的设置(i)将模型中的因果机制与真实数据生成过程的因果结构分离开来,(ii)通过考虑潜在共同原因的推断,更倾向于自然语言预测,(iii)探讨了潜在共同原因的贝叶斯证据积累是否以及如何在允许跨上下文泛化到新测试用例的表征和机制中实现。

英文摘要

We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the model from the causal structure of the true data-generating process, (ii) orients more towards natural language prediction by considering inference of latent common causes, and (iii) considers whether and how Bayesian evidence accumulation for latent common causes can be implemented in representations and mechanisms that allow for cross-context generalization to novel test cases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑