arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12935cs.AIcs.LG

扰动响应中证据、矛盾与脆弱性的分解

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

Lei You

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对扰动方法仅能反映模型响应程度的问题,提出DECAF分解方法将响应分为证据、矛盾、脆弱性分量,在多数据集实验中验证其准确性,且效率优于通用归因基线。

中文摘要 AI 辅助

扰动方法通过测量输入改变下的预测变化来解释模型决策,但响应幅度仅能说明模型的反应程度,无法表明该反应的含义。相同的幅度可能支持最终事实-反事实差异、反对该差异,或在扰动路径中强烈出现却在终点消失。因此,我们跟踪成对输入逐步揭示时对比的发展情况,利用最终对比来解释轨迹。我们引入DECAF(证据、矛盾与脆弱性的分解,Decomposition of Evidence, Contradiction, And Fragility),其将对齐、反对和终点-无效响应分别归为证据E、矛盾C和脆弱性F。该分解精确保留普通幅度,满足Abs = E + C + F,且在终点相关公理下是唯一的。在受控的视觉和表格设置中,三个分量与独立测量的行为跟踪一致。在对72个模型的ImageNet-9审计中,我们比较响应幅度几乎相同但独立测量行为不同的情况,最大的DECAF分量在96.4%的案例中与观察到的行为一致,而仅幅度的一致性为35.0%。仅改变揭示路径会使总响应增加近80%,但证据几乎不变,而脆弱性增长超过4倍。在FunnyBirds和ImageNet-1k上,仅前向的短DECAF轨迹优于测试的通用归因基线。在1B规模的DINOv2模型上,短轨迹与强梯度基线匹配,却具有低4.75倍的墙钟时间和低2.36倍的峰值内存。

英文摘要

Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.

发表机构

  • Technical University of Denmark(丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑