发表机构
Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出有符号曲率准则,证明一次性反事实步骤的有效性取决于路径曲率非负,并给出基于曲率界的单查询规则及校准方法,实验验证其有效性。
AI 中文摘要
闭式反事实解释将被拒绝的用户沿分类器得分 f 的单位梯度 ĝ 移动承诺距离 d_p=|f(x)|/‖∇f(x)‖,在该距离处线性化得分达到零。我们研究此一次性步骤何时成功以及额外的模型查询会带来什么变化。在一阶近似下,当路径曲率 κ=ĝ^T∇²f(x)ĝ 非负时,该步骤恰好终止于有利侧。在80个浅层模型中,被拒绝用户中步骤终止于有利侧的比例与 κ≥0 的比例之间的相关系数为 r=0.985,尽管在 Fashion-MNIST 上前者平均比后者低8.2个百分点。任何仅使用得分值和梯度的规则,对于路径曲率以 K 为界的任意得分,都不可能在不以 Kd_p²/‖∇f(x)‖ 量级过度冲过某些点的情况下保证有效性。当曲率还满足 Lipschitz 条件且步骤较短时,在承诺点处对 f 的一次评估在确定性单查询规则中达到极小极大速率,这些规则已知曲率界及其 Lipschitz 常数,并且分裂保形校准使得此类规则以至少 1-δ 的概率达到首次穿越或弃权(不执行)。使用非对称曲率惩罚训练,在易欠冲的浅层数据(Fashion-MNIST、COMPAS)上,99-100% 的路径在承诺步骤内穿越,其过冲量约为对称惩罚的4-22倍。由于 κ 和 d_p 依赖于得分的缩放方式,部分增益可能来自更长的承诺步骤,并且在匹配有效性的情况下,对短暂训练模型的较小审计未发现相对于调整膨胀的均匀优势。在沿射线进行逐用户线搜索可行的情况下,该方法在网格分辨率下是精确的,并且更可取。
英文摘要
Closed-form recourse moves a rejected user along the unit gradient $\hat g$ of the classifier score $f$ by the promised distance $d_p=|f(x)|/\|\nabla f(x)\|$, at which the linearized score reaches zero. We ask when this one-shot step succeeds and what additional model queries change. To leading order the step ends on the favorable side exactly when the path curvature $κ=\hat g^\top\nabla^2 f(x)\,\hat g$ is nonnegative. Across 80 shallow models, the fraction of rejected users whose step ends there and the fraction with $κ\ge0$ correlate at $r=0.985$, although on Fashion-MNIST the first falls below the second by 8.2 points on average. No rule that uses only the score value and gradient can be valid for every score with path curvature bounded by $K$ without overshooting some by order $Kd_p^2/\|\nabla f(x)\|$. When the curvature is also Lipschitz and the step is short, one evaluation of $f$ at the promised point attains the minimax rate among deterministic one-query rules that know the curvature bound and its Lipschitz constant, and split-conformal calibration makes such a rule reach the first crossing or abstain with probability at least $1-δ$. Training with an asymmetric curvature penalty lets 99-100% of paths cross within the promised step on undershoot-prone shallow data, at about 4-22 times the overshoot of symmetric penalties (Fashion-MNIST, COMPAS). Because $κ$ and $d_p$ depend on how the score is scaled, part of this gain can be a longer promised step, and at matched validity a smaller audit of briefly trained models finds no uniform advantage over tuned inflation. Where a per-user line search along the ray is affordable, it is exact to grid resolution and preferable.
Commentsv2: minor corrections and clarifications. Main text 8 pages; 58 pages including appendix and checklist