arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08869cs.CL

序数分类中的位置偏差:系统评估

Are LLMs Positionally Consistent Ordinal Classifiers? A Systematic Evaluation

发表机构康奈尔大学 · 华盛顿大学 · 香港中文大学(深圳)
另 1 家 · 查看机构详情
  • Cornell University(康奈尔大学)
  • University of Washington(华盛顿大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • University of California, Davis(加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

Yu Wang, Zhe Zhou, Menglin Liu, Ge Shi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究系统评估了序数分类中大型语言模型的位置偏差,发现其普遍存在且现有校正方法效果有限,提出需结合预测性能与稳定性选择序数分类系统。

中文摘要 AI 辅助

大型语言模型正越来越多地被用于序数分类,但提示组织的语义等价改变会改变它们的预测结果。我们开展系统实验,以表征来自标签顺序、示例顺序和示例位置的位置偏差。首先,我们在一项通用序数分类任务上对十个前沿大型语言模型(LLM)应用三种探测方法,所有模型都对这三种位置来源敏感,表明该问题普遍存在。其次,我们在五个数据集上改变八个提示级、任务级和模型级因素;准确率和稳定性经常不一致,仅较低的基数规模能始终同时提升两者。第三,我们比较点式、成对和列表式推理、替代聚合与去偏方法,以及联合配置;所测试的校正方法无法提供可靠补救,而基于比较的列表式公式能提供最佳平衡,但在不同模型和偏差来源间的迁移效果不均。这些发现表明,位置鲁棒性取决于完整系统配置而非仅模型本身,因此应同时根据预测性能和稳定性来选择序数分类系统。

英文摘要

Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from label order, demonstration order, and demonstration placement. First, we apply the three probes to ten frontier LLMs on a common ordinal-classification task; every model is sensitive to all three positional sources, showing that the problem is pervasive. Second, we vary eight prompt-, task-, and model-level factors across five datasets; accuracy and stability are often misaligned, and only lower scale cardinality consistently improves both. Third, we compare pointwise, pairwise, and listwise inference, alternative aggregation and debiasing methods, and joint configurations; the tested corrections do not provide a reliable remedy, while a comparison-based listwise formulation offers the best balance but transfers unevenly across models and bias sources. These findings show that positional robustness depends on the full system configuration rather than the model alone. Ordinal-classification systems should therefore be selected jointly for predictive performance and stability.

↑