更多选择,更少决策:JEV类直接决策模型中的序数尺度偏差
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models
浏览论文内容
中文总结 AI 辅助
研究发现JEV类直接决策模型存在序数尺度利用偏差:在36个序数数据集上,模型仅使用67%-76%的有效金标签支持,且随尺度K增大利用率下降至26%-75%;通过BA-LoRA后训练可将利用率从约47%提升至86%,表明该压缩是可习得和可修改的。
中文摘要 AI 辅助
直接决策模型将文本转化为低延迟的结构化标签和分数,使其在分类和自动评估中具有吸引力。然而,可靠性不仅仅需要准确性:模型还必须忠实地使用用户提供的序数决策尺度。我们分析了JEV 1.13和三个开源的KEV模型。我们的研究从ANLI开始,在该数据集中,JEV将38.8%的所有预测和51.3%的错误分配给“中性”类别,尽管准确率达到74.95%,金标签几乎平衡,且候选位置也平衡。在36个序数数据集中,最终决策仅使用了有效金标签支持的67%至76%,而在四个名义任务中这一比例为87%至102%。随机化候选顺序会削弱但不会消除这种压缩。在保持项目和源分数不变的同时平衡金标签支持和位置,我们将尺度从K=2细化到K=14;每个模型的利用率都会下降,在K=14时达到26%至75%,尽管大多数模型的候选概率仍然很宽。针对性的BA-LoRA后训练在两种KEV规模下,将八个监督尺度上的金标签相对利用率从大约47%提高到86%,表明这种压缩是习得的且可修改的,而非不可改变的结构限制。我们将此称为序数尺度利用偏差:决策阶段的候选空间压缩,这不同于准确性、金标签不平衡、固定位置或候选数量单独造成的影响。代码和数据可在以下网址获取:https://this-url
英文摘要
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8\% of all predictions and 51.3\% of errors to Neutral despite 74.95\% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76\% of the effective gold support, versus 87--102\% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from $K=2$ to $14$; utilization falls for every model and reaches 26--75\% at $K=14$, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47\% to 86\% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia
发表机构
- Jilin University(吉林大学)
机构由 AI 辅助整理,请以论文原文为准。