arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

决策偏移、标签功能丧失与正确性门控多教师蒸馏中不确定的接地审计

Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation

Xiaofei Feng

arXiv 2609.09702首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在固定实验中考察正确性门控多教师蒸馏,发现加权策略虽改变决策分布但丧失标签功能,且人类审计无法证明接地增益,相对硬过滤无增量决策收益。

AI 中文摘要

候选决策的正确性与理由接地是不同的目标。我们在一个固定实验中考察了正确性门控的多教师蒸馏。八个实验组共享4,330个来源、一个63.9M参数的student模型、12,990个优化行、406次更新、证据输入和一个解码器;七个基于教师的实验组使用一个固定的三响应池。在267个保留样本上评估了三个种子。相对于未过滤的蒸馏,正确性加权实验组在准确率上相差+0.1660(95%观测矩阵区间[0.0670, 0.2455]),在五标签宏F1上相差+0.1323([0.0916, 0.1731]),在任务定义的条件不安全动作率上相差-0.4979([-0.5926, -0.3686])。这些偏移并不暗示一致更好的行为。源标签SFT具有最高的平均宏F1(0.586)。加权实验组在每个种子中的Refuted召回率均为零,且两个种子将所有167个声明示例分配为NotEnoughInfo。在一个参考种子上的可用性修正审计中,加权和未过滤输出分别有0/20和1/20的证据支持正例,以及20/20和19/20的包含不支持材料的正例。样本未配对,来源重叠未序列化,修正发生在自动摘要之后但在标注之前。因此,该审计无法估计共同来源的接地效应,并且对于系统层面的改进或损害不确定。硬过滤已经达到0.660的准确率、0.530的宏F1和0.135的条件不安全率。实现的加权实验组相对于硬过滤未显示出已证明的增量决策收益。这个固定矩阵失败分析显示了带有标签功能丧失的决策重新分配;可用的人类审计并未确立接地增益。

英文摘要

Candidate decision correctness and rationale grounding are different objectives. We examine correctness-gated multi-teacher distillation in a fixed experiment. Eight arms share 4,330 sources, a 63.9M-parameter student, 12,990 optimization rows, 406 updates, evidence inputs, and a decoder; seven teacher-based arms use one fixed three-response pool. Three seeds are evaluated on 267 held-out examples. Relative to unfiltered distillation, the correctness-weighted arm differed in accuracy by +0.1660 (95% observed-matrix interval [0.0670, 0.2455]), five-label macro-F1 by +0.1323 ([0.0916, 0.1731]), and task-defined conditional unsafe-action rate by -0.4979 ([-0.5926, -0.3686]). These shifts do not imply uniformly better behavior. Source-label SFT had the highest mean macro-F1 (0.586). The weighted arm had zero Refuted recall in every seed, and two seeds assigned NotEnoughInfo to all 167 claim examples. In an availability-amended audit at one reference seed, weighted and unfiltered outputs had 0/20 versus 1/20 evidence-supported positives and 20/20 versus 19/20 positives containing unsupported material. Samples were non-paired, source overlap was not serialized, and the amendment followed automatic summarization but preceded annotation. The audit therefore cannot estimate a common-source grounding effect and is inconclusive about system-level improvement or harm. Hard filtering already achieved 0.660 accuracy, 0.530 macro-F1, and 0.135 conditional unsafe rate. The implemented weighted arm showed no demonstrated incremental decision benefit over hard filtering. This fixed-matrix failure analysis shows decision redistribution with lost label functionality; the available human audit does not establish a grounding gain.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑