发表机构
Amity Research and Application Center (ARAC), Amity Group(Amity研究与应用中心(ARAC),Amity集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
实证研究比较行为语言模型中评分式与生成式读出,发现评分式在多数任务中排序更准,建议保留生成推理但用评分头排序。
AI 中文摘要
针对客户行为进行微调的语言模型能够预测结果并生成解释,但这些读出(readouts)常常被视为可互换的。在保持模型检查点和提示内容固定的情况下,我们比较了通过对答案令牌进行评分获得的概率与在书面推理之后生成的预测。在涵盖三个市场中四项零售任务的13个模型-领域单元中,包括两个使用完全公开数据和检查点的单元,评分式读出在13个单元中的12个中更准确地排列结果(双侧符号检验,p约0.003),在接收者操作特征曲线下面积(AUC)上高出1.5至14.5个百分点。配对自助法置信区间在每个新测量的单元中均排除零。该差距随任务特定监督以及训练与服务格式之间的不匹配而变化,范围从未调优基础模型的-2.2个百分点到推理格式监督的+13.7个百分点。对约9,000条推理的分析确定了两个相关因素:对主导预测特征的依赖减少以及向固定表述的收敛。概率饱和并不追踪该差距。第三种读出,在任何判定之前引出概率,改善了校准(Brier分数从0.47降至0.15),同时排序在评分的噪声范围内,但仅适用于训练中表示的结果率;当评分头已校准时,其表现不如评分。我们通过每种读出所匹配的目标来解释这些差异,识别出缩小差距的训练选择,并建议保留生成的推理,同时从评分头获取排序。
英文摘要
Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with predictions generated after a written rationale. Across 13 model-domain cells covering four retail tasks in three markets, including two using fully public data and checkpoints, the scored readout ranks outcomes more accurately in 12 of 13 cells (two-sided sign test, p approximately 0.003), by 1.5 to 14.5 points in area under the receiver operating characteristic curve (AUC). Paired bootstrap confidence intervals exclude zero in every newly measured cell. The gap varies with task-specific supervision and mismatch between training and serving formats, ranging from -2.2 points for an untuned base model to +13.7 for rationale-format supervision. Analysis of approximately 9,000 rationales identifies two correlates: reduced reliance on the dominant predictive feature and convergence on stock formulations. Probability saturation does not track the gap. A third readout, eliciting a probability before any verdict, improves calibration (Brier score from 0.47 to 0.15) while ranking within noise of scoring, but only for outcome rates represented in training; it is worse than scoring when the scored head is already calibrated. We interpret these differences through the objectives matched by each readout, identify training choices that narrow the gap, and propose retaining generated rationales while sourcing ranking from the scored head.
Comments12 pages, 1 figure, 2 tables