PACT:面向单标记类型化决策的成对锚定校准调优
PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions
浏览论文内容
中文总结 AI 辅助
PACT利用对比对结构,通过四个训练项和上下文温度,在不需新标注下提升单标记类型化决策模型的稳健性与稳定性,降低位置偏差和序数误差,而非提高原始准确率。
中文摘要 AI 辅助
单标记类型化决策模型通过读取单个位置上几个单字母答案代码的逻辑值来回答模式问题:它们速度快,并为每个允许的答案返回概率,但它们是使用普通交叉熵训练的,忽略了训练数据中的大部分结构。我们研究了这样一个模型,其数据被整理为对比对——两个上下文在一个编辑过的翻转答案的事实上有所不同——每个对比对都带有机器可验证的证书,证明删除决定性句子会使该事实变为未知。我们提出了PACT,它将这种结构转化为四个训练项,无需新的标注:一个对每对共享逻辑偏移不变的差分差分边际项,一个针对答案代码位置偏差的排列一致性项,一个基于证书验证的消融上下文的证据必要性项,以及一个用于评分字段的序数传输成本,外加一个三参数上下文温度。在冻结的324项保留集上,使用三个随机种子,PACT在准确率上与已发表配方相当(84.6%对比85.2%;每个种子的McNemar p≥0.50),同时在所有运行中给出最低的位置偏差(重新标记下答案翻转9.8%对比13.8%)和评分字段上最低的序数误差(MAE 0.232对比0.311)。与使用相同优化器和调度但仅使用交叉熵的对照组相比,PACT在三个种子中的两个上显著更准确,将种子间差异减半,并将NLL降低26%。种子匹配的消融和预先指定的证伪测试精确定位了这些收益:没有单个项能提高原始准确率,该方法的价值在于稳健性和稳定性,而非标题性的准确率。代码、数据划分、训练好的适配器以及所有运行记录均可在该https URL获取。
英文摘要
Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores most of the structure in their training data. We study such a model whose data is curated as contrastive pairs---two contexts that differ in one edited fact that flips the answer---each carrying a machine-checked certificate that deleting the decisive sentence makes the fact unknown. We propose PACT, which turns this structure into four training terms that need no new annotation: a difference-in-differences margin over each pair that is invariant to any shared logit offset, a permutation-consistency term against answer-code position bias, an evidence-necessity term on certificate-verified ablated contexts, and an ordinal transport cost for rubric fields, plus a three-parameter contextual temperature. On a frozen 324-item holdout with three seeds, PACT matches the published recipe in accuracy ($84.6\%$ vs. $85.2\%$; McNemar $p \ge 0.50$ at every seed) while giving the lowest position bias of all runs (answer flips under relabelling $9.8\%$ vs. $13.8\%$) and the lowest ordinal error on rubric fields (MAE $0.232$ vs. $0.311$). Against a control with the same optimiser and schedule but cross-entropy only, PACT is significantly more accurate at two of three seeds, halves the seed-to-seed spread and lowers NLL by $26\%$. Seed-matched ablations and pre-specified falsification tests locate these gains precisely: no single term raises raw accuracy, and the method's value lies in robustness and stability rather than headline accuracy. Code, data splits, trained adapters, and all run records are available at https://github.com/BennyLinntu/PACT-Pairwise-Anchored-Calibrated-Tuning-for-Single-Token-Typed-Decisions.
发表机构
- Victoria University of Wellington(惠灵顿维多利亚大学)
机构由 AI 辅助整理,请以论文原文为准。