arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体决策的因果归因:估计量、耦合与可追溯性规范

Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification

Ajay Pravin Mahale

arXiv 2609.06445首次发表:更新:

AI 中文总结

针对智能体决策的因果归因,提出估计量框架并揭示边际与公共随机数总效应在特定条件下的失效,推导耦合保持直接效应可估,并给出可追溯性规范以应对法规时间差。

AI 中文摘要

高风险AI系统的提供者必须保存使决策可追溯的记录,而对于智能体系统,尚未确定这些记录必须包含什么内容才能进行事后因果归因。我们给出了估计量框架,然后给出了其失效的条件。我们将先前工作测量的边际总效应与隔离某一步骤自身贡献的公共随机数总效应区分开来,在固定下游条件下增加了自然直接效应,并将估计量与手工推导进行核对。然后,两个估计目标都以相同的方向失效。在边际估计量下,一个因果惰性步骤与决定性步骤在我们的植入链的每次运行中具有相同的总效应,这是一个代数恒等式,而非单次抽取的巧合。在公共随机数下,当执行步骤翻转时,决定性步骤在约十分之一的运行中恰好返回零,而其直接效应为0.25,且它明显起作用;精确的零并不证明某一步骤未做任何事情,我们在此说明这一点,而非放在局限性中。我们推导了在上下文发散后保持直接效应可估计的耦合,并给出了其退化的闭式解,并表明建立在自然排序上的中介份额在抑制下不是份额:在直接路径和中介路径相反的地方,它超过1,并将被抑制的组件排在纯中介之上。我们发布了差异实验的预注册,而不是结果,因为该实验所需的实时管道在研究窗口内不可用。我们贡献了此类备案所需的可追溯性规范,以应对该法案日历所开启的空白:第86条的解释权自2026年8月2日起适用,而根据(EU) 2026/1744号法规,第12条的日志记录和附件IV的文档编制被推迟至2027年12月2日。

英文摘要

A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditions under which it fails. We separate the marginal total effect that prior work measures from a common-random-numbers total effect that isolates a step's own contribution, add the natural direct effect under a pinned downstream, and check the estimators against hand derivations. Both estimands then fail, in the same direction. Under the marginal estimand a causally inert step has the identical total effect to the decisive one on every run of our planted chain, an algebraic identity and not a coincidence at one draw. Under common random numbers the decisive step returns exactly zero on the runs where the executing step flips, about one in ten, while its direct effect there is 0.25 and it demonstrably acts; an exact zero does not certify that a step did nothing, and we put that here rather than in the limitations. We derive the coupling that keeps the direct effect estimable once contexts diverge, with a closed form for its degradation, and show that the mediated share on which a natural ranking is built is not a share under suppression: where the direct and mediated paths oppose, it exceeds one and ranks a suppressed component above a pure mediator. We publish the discrepancy experiment's pre-registration rather than a result, because the live pipeline it requires was not available in the study window. We contribute the traceability specification such a filing would need, against a gap the Act's calendar opens: Article 86's right to an explanation has applied since 2 August 2026, while the Article 12 logging and Annex IV documentation that could evidence one were deferred to 2 December 2027 by Regulation (EU) 2026/1744.

Comments23 pages, 2 figures, 2 tables, 3 propositions with proofs. Pre-registered discrepancy experiment published in full and not run; no empirical result is claimed. Companion to arXiv:2608.13754. Code, validators, derivations and the traceability specification: https://github.com/designer-coderajay/Causal-Attribution-for-Agentic-Decisions

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑