arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

执行因果结构学习与线性注意力变换器

Executing Causal Structure Learning with Linear-Attention Transformers

Amartya Roy, Sayar Karmakar

arXiv 2610.10395首次发表:更新:

发表机构

Indian Institute of Technology Delhi; University of Florida(印度理工学院德里分校; 佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构造固定权重线性注意力变换器精确执行因果结构学习算法,证明乘子保留关键性,实验验证精度与迁移局限,区分算法执行与因果恢复。

AI 中文摘要

变换器能够对其输入中给出的数据执行算法。我们询问它们是否也能对因果发现执行同样的操作。我们研究了一种标准的连续方法,该方法在强制无环性的同时反复更新候选因果图。我们显式地构造了一个固定权重的变换器,其前向传播恰好重现了该方法的一次更新,因此重复的块可以重现其优化轨迹。该变换器在更新之间携带当前图以及算法的乘子。我们表明,保留乘子对于精确执行至关重要,因为不同的乘子值可能导致不同的下一次更新。我们还给出了条件,在这些条件下,在固定阶段内,达到目标精度所需的更新次数可以预先计算,并且随着深度增加,舍入误差保持有界。实验表明,所构造的块与参考更新在浮点精度上一致,而在合成数据和七个已发布的基准网络拓扑上的算术重放继承了参考求解器的成功和失败。这将准确的算法执行与准确的因果恢复区分开来。相比之下,在我们训练预算下测试的普通注意力模型不能可靠地执行更新或迁移到更大的图。梯度训练是否能在该构造的架构类别中学习执行器仍然是一个开放问题。

英文摘要

Transformers can execute algorithms on data given in their input. We ask whether they can do the same for causal discovery. We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity. We explicitly construct a fixed-weight transformer whose forward pass exactly reproduces one update of this method, so repeated blocks reproduce its optimization trajectory. The transformer carries the current graph and the algorithm's multiplier between updates. We show that retaining the multiplier is essential for exact execution, since different multiplier values can lead to different next updates. We also give conditions under which, within a fixed stage, the number of updates needed to reach a target accuracy can be computed in advance and rounding errors stay bounded as depth grows. Experiments show that the constructed block agrees with a reference update to floating-point precision, while arithmetic replay on synthetic data and seven published benchmark network topologies inherits the reference solver's successes and failures. This separates accurate algorithm execution from accurate causal recovery. In contrast, the ordinary attention models tested under our training budgets do not reliably execute the update or transfer to larger graphs. Whether gradient training can learn an executor in the architecture class of the construction remains open.

Comments32 pages, 8 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑