AI 中文总结
本研究通过精确分解和对照实验证明,列表式重排序中有限视图不稳定性得分对未见排列的预测能力依赖于验证目标与探针观测不相交,重用视图会高估泛化性。
AI 中文摘要
列表式语言模型重排序器在等价候选排列之间常常产生不一致的结果。因此,有限不稳定性诊断被用于推动额外的采样、聚合或选择性计算。然而,与重用探针视图的验证统计量的关联,并不一定能分离出关于未见排列的预测信息。共享测量可能引发经典的“部分-整体”关联。我们研究了这一现象如何影响关于“有限视图不稳定性得分能预测未见排列”的主张。我们推导了精确的有限视图分解,并前瞻性地比较了零个、一个和两个重用视图,包括一个完全不相交的四视图目标。该研究涵盖两个固定的70亿参数模型家族和两个推荐数据集,使用受控列表进行带符号的离线分析,并使用未经处理的检索器列表进行无目标复制。在四个受控块上,完全不相交的相关性较弱或异质(-0.061至0.281),而重用两个探针视图则产生0.600至0.718的相关性;所有配对对比差异均较大(0.436至0.661)且经Holm校正后显著。重叠效应在所有四个未经处理的列表块中均为正。将探针从两个视图增加到四个视图,仅在其中一个块中明显改善了不相交可靠性。此外,探针能预测聚合移动幅度(Spearman rho = 0.142至0.426),但不能预测稳定的带符号目标收益,并且12个固定比例探针路由点中有7个在测量成本下被严格支配。因此,当目标估计量是关于未见扰动行为的预测信息时,验证目标必须与探针在观测上不相交,以分离该信息;带符号效用和成本敏感决策仍是独立的问题。
英文摘要
Listwise language-model rerankers often disagree across equivalent candidate permutations. Finite instability diagnostics are therefore used to motivate additional sampling, aggregation, or selective computation. But an association with a validation statistic that reuses the probe views need not isolate predictive information about unseen permutations. Shared measurements can induce classical part-whole association. We study how this affects claims that a finite-view instability score predicts unseen permutations. We derive the exact finite-view decomposition and prospectively compare zero, one, and two reused views, including a fully disjoint four-view target. The study covers two pinned 7B model families and two recommendation datasets, with controlled lists for signed offline analysis and untouched retriever lists for target-free replication. On the four controlled blocks, fully disjoint correlations are weak or heterogeneous (-0.061 to 0.281), whereas reusing both probe views yields 0.600 to 0.718; all paired contrasts are large (0.436 to 0.661) and Holm-significant. The overlap effect is positive in all four untouched-list blocks. Increasing the probe from two to four views clearly improves disjoint reliability in only one block. Moreover, the probe predicts aggregation-movement magnitude (Spearman rho = 0.142 to 0.426) but not stable signed target benefit, and 7 of 12 fixed-fraction probe-routing points are strictly dominated at measured cost. Thus, when the intended estimand is predictive information about unseen perturbation behavior, validation targets must be observation-disjoint from the probe to isolate that information; signed utility and cost-sensitive decisions remain separate questions.
Comments8 pages, 2 figures, 4 tables