当编辑定位放大相对选择偏差:梯度几何、目标不匹配与重要性加权
When Edit Localization Amplifies Relative Selection Bias: Gradient Geometry, Target Mismatch, and Importance Weighting
浏览论文内容
中文总结 AI 辅助
本研究分析编辑定位对相对选择偏差的影响,提出梯度分解与重要性加权方法,通过蒙特卡洛模拟和翻译实验证明硬局部化在相对偏差上更高但绝对偏差更低。
中文摘要 AI 辅助
人类纠错可识别可编辑的跨度,但接收纠错的示例可能来自选择性反馈通道。我们在固定模型检查点处通过将局部化梯度分解为已编辑和保留的未触碰分量来分析这一交互。平方相对选择偏差是二次型之比,其导数具有显式二次多项式的符号。局部化可以增加、减少或非单调地改变该诊断指标;其方向取决于分量偏差和几何结构。Oracle重要性加权为每个固定局部化目标恢复总体均值,但这些目标具有不同的目标。针对一个常见的全梯度目标,我们推导了有限样本均方误差、解析最优保留系数以及固定裁剪扩展。精确的有限总体计算和每个样本量10,000次蒙特卡洛重复验证了恒等式和反例。公开的人类编辑后实验使用两个翻译方向和预训练模型,并声明了合成选择。英德扩展区分了7389万个原生参数。在所有三个声明设置中,硬局部化具有更高的相对偏差但更低的绝对偏差,相比完全保留。未触碰分量偏差非零,且所选分量均值具有负内积,因此一般准则适用于简单无偏/对齐解释失败的情况。输出偏差诊断显示两个方向上相对偏差的相同端点排序,并有一个内部最大值。机制重用每种语言的记录并包含混合;它们不是独立复制。证据将相对放大与绝对梯度误差分开,并建立了估计性质,而不推断翻译质量提升或识别实际投诉倾向。
英文摘要
Human corrections identify editable spans, but the examples receiving corrections may come from a selective feedback channel. We analyze this interaction at a fixed model checkpoint by decomposing a localized gradient into edited and retained untouched components. Squared relative selection bias is a ratio of quadratics whose derivative has the sign of an explicit quadratic polynomial. Localization can increase, decrease, or nonmonotonically change this diagnostic; its direction depends on component biases and geometry. Oracle importance weighting recovers the population mean for each fixed localization objective, but these objectives have different targets. Against one common full-gradient target, we derive the finite-sample mean-squared error, an analytic optimal retention coefficient, and a fixed-clipping extension. Exact finite-population calculations and 10,000 Monte Carlo repetitions per sample size verify the identities and counterexamples. Public human-post-edit experiments use two translation directions and pretrained models, with declared synthetic selection. An English-German extension differentiates 73.89 million native parameters. In all three declared settings, hard localization has higher relative bias but lower absolute bias than full retention. Untouched-component biases are nonzero and selected component means have negative inner products, so the general criterion applies where the simple unbiased/aligned explanation fails. Output-bias diagnostics show the same endpoint ordering of relative bias across both directions, with one interior maximum. Mechanisms reuse each language's records and include a mixture; they are not independent replications. The evidence separates relative amplification from absolute gradient error and establishes estimation properties, without inferring translation-quality gains or identifying actual complaint propensities.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。