arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01108cs.LG

复现TRACE:从业者指南——其阈值与粒子预算

Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget

Alex Chadyuk, Alicia Zhang, Roy Kucukates

首次发表
浏览论文内容

中文总结 AI 辅助

本文复现TRACE算法,发现其最优阈值与真值边际相关,单一全局阈值对滞后1边效果好、对高阶边差,基准滞后衰减会混淆算法极限,粒子数2时F1饱和,并提炼出五条从业者规则。

中文摘要 AI 辅助

TRACE(Math和Lienhart,arXiv:2602.01135)通过在固定τ阈值下对每个位置的条件互信息估计值进行阈值化处理,从预训练自回归序列模型中读取事件类型的因果图。我们独立复现了其核心合成结果:在验证集上选择τ时,词汇量为1000时,针对精确干预真值的每序列平均F1值达到0.90-0.91(论文结果为0.91),词汇量从100到2000时,该值为0.86-0.91。首先,最优阈值被固定在真值边际,而非任何常数:在每个规模下,τ*处的误差跨越定义真值的Δ=0.05边际(遗漏的真实边恰好位于该边际上方,被接受的假边恰好位于下方),且盲最优值接近Δ/2乘以估计器的校准值,在5000个样本的外测试中得到验证。其次,在单一全局阈值下,TRACE主要恢复直接的相邻影响图:滞后1的真实边召回率为0.97-0.99,而滞后2或更久的真实边召回率低几个数量级——这是在真值未知时,精确直接因果效应检验需要随机中介位置所带来的读取规模代价。按滞后划分的阈值系列可恢复三分之一到一半的滞后2真值;在滞后均匀数据上,一个经验证的阈值对各滞后的召回率为0.40-0.87,在滞后3-6时比原子干预对照低8-26个百分点。第三,论文合成基准的默认滞后衰减将约85%的干预真值集中在滞后1,其余部分推至估计器的噪声底以下,因此核心F1值仅证明滞后1的恢复,并将基准的偏斜与算法自身的极限混为一谈;更平缓的衰减可将两者区分开。第四,在选定阈值下,F1从N=2个粒子开始达到饱和——这是阈值相对于噪声底的边际特性,而非估计器的特性,估计器随N^(-1/2)收敛。我们提炼出五条从业者规则。

英文摘要

TRACE (Math & Lienhart, arXiv:2602.01135) reads causal graphs over event types out of a pretrained autoregressive sequence model by thresholding a per-position conditional-mutual-information estimate at a fixed tau. We independently replicate its headline synthetic result: with tau selected on a validation split, mean per-sequence F1 against exact interventional truth reaches 0.90-0.91 at vocabulary size 1000 (paper: 0.91) and 0.86-0.91 from 100 to 2000. First, the optimal threshold is pinned to the truth margin, not to any constant: at every size the errors at tau* straddle the delta = 0.05 margin defining ground truth (missed true edges lie just above it, accepted false ones just below), and the blind optimum lands near delta/2 times the estimator's calibration, confirmed out of sample at 5000. Second, at a single global threshold TRACE mostly recovers a direct, adjacent-influence graph: lag-1 true edges are recalled at 0.97-0.99, while true edges at lag 2 or more read orders of magnitude lower---the reading-scale price of randomizing mediating positions, which an exact test of direct causal effect requires when the truth is unknown. A per-lag threshold family recovers a third to a half of lag-2 truth; on lag-uniform data one validated threshold recalls every lag at 0.40-0.87, 8-26 pp below an atomic-intervention control at lags 3-6. Third, the default lag decay of the paper's synthetic benchmark concentrates about 85% of interventional truth at lag 1 and pushes the rest below the estimator's noise floor, so headline F1 there certifies lag-1 recovery only and conflates the benchmark's skew with the algorithm's own limit; a flatter decay separates the two. Fourth, F1 saturates from N = 2 particles at the selected threshold---a property of the threshold's margin over the noise floor, not of the estimator, which converges as N^(-1/2). We distill five practitioner rules.

发表机构

  • LotusFlare Inc.(莲花flare公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑