arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越极化:点式重排序中思维链的生成约束

Beyond Polarization: The Generative Constraint of Chain-of-Thought in Pointwise Reranking

Xiaoyang Chen, Jie Liu, Haijin Liang, Haibo Shi, Jin Ma, Ben He, Yingfei Sun, Dezhi Ye

arXiv 2608.30398首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Tencent(中国科学院大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对点式重排序中思维链模型性能弱于直接评分模型的问题,通过实证研究及多种干预措施,发现排序差距源于离散文本传递连续相关性语义的瓶颈,而非易解决的训练偏差。

AI 中文摘要

在点式文档重排序任务中,思维链(Chain-of-Thought)模型通常表现不如直接评分模型。现有诊断研究将此归因于分类性能较差、分数极化或校准失效,但针对性训练能否缩小该差距仍不明确。本实证研究首先确认,该差距在参数规模达32B时仍保持稳定,排除了模型与数据容量混淆变量的影响。随后,本研究运用强化学习、细粒度监督及架构解耦等压力测试,明确修复这些偏差。尽管这些干预措施提升了分类准确率与绝对分数,但相对排序差距依然存在。这些发现表明,在点式评分范式内,通过离散文本传递连续相关性语义会限制排序信号的分辨率,揭示出当前标准方法下稳定且难以克服的瓶颈,而非易解决的训练偏差。

英文摘要

In pointwise document reranking, Chain-of-Thought models typically underperform direct scoring models. While existing diagnostics attribute this to inferior classification, score polarization, or calibration breakdown, whether targeted training can bridge this gap remains unclear. Our empirical study first confirms that this gap is stable across scales up to 32B parameters, ruling out model and data capacity confounders. We then apply stress tests utilizing reinforcement learning, fine-grained supervision, and architectural decoupling to explicitly repair these deviations. Although these interventions improve classification accuracy and absolute scores, the relative ranking gap persists. These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution, revealing a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.

CommentsAccepted at EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑