arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cloud-ScPO:面向大语言模型推理的半监督偏好优化的隐状态几何

Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning

Yuzhou Liu, Xiyang Hu

arXiv 2608.01014首次发表:更新:

AI 中文总结

本研究提出Cloud-ScPO框架,利用少量标注集构建参考云,结合拓扑引导的偏好挖掘与自一致性,在GSM8K等数据集上较ScPO提升LLM数学推理性能。

AI 中文摘要

偏好优化可提升大语言模型(LLM)的数学推理能力,但可靠的“选择-拒绝”对通常需要经验证的答案、人工标注或外部奖励模型。本研究探究半监督场景下,偏好监督是否可源自模型的内部表示几何。分析显示,不同数学问题生成的推理轨迹会形成结构化的全局点云,其中正确与错误轨迹呈现不同的几何组织。基于此,提出拓扑引导的偏好挖掘框架Cloud-ScPO,利用少量标注集构建多个正确与错误参考云;每条轨迹由平均池化的隐状态表示,通过跨参考库平均的组件级软k近邻度量,按连接诱导的组件评分。将跨问题的云信号与提示级自一致性结合:自一致性确定答案级偏好方向,云评分选择具体轨迹并按评分间隔筛选对。在GSM8K和MATH-Numeric的4种模型设置下实验表明,Cloud-ScPO较ScPO持续提升,在GSM8K上最高提升4.49%,在MATH-Numeric上最高提升4.19%;成对分析进一步显示,Cloud-ScPO在保持相当正确性可靠性的同时,更有效地将有价值的选择轨迹与不完整、重复或其他低质量的拒绝响应区分开。

英文摘要

Preference optimization improves mathematical reasoning in large language models (LLMs), but reliable chosen-rejected pairs usually require verified answers, human annotations, or external reward models. We investigate whether preference supervision can instead be derived from the model's internal representation geometry in a semi-supervised setting. Our analysis shows that reasoning trajectories generated across different mathematical problems form structured global point clouds in which correct and incorrect trajectories exhibit different geometric organization. Based on this observation, we propose Cloud-ScPO, a topology-guided preference-mining framework that uses a small labeled set to construct multiple correct and incorrect reference Clouds. Each trajectory is represented by a mean-pooled hidden state and scored against connectivity-induced components using a component-level soft $k$-nearest-neighbor measure averaged across reference banks. We combine this cross-problem Cloud signal with prompt-level self-consistency: self-consistency determines the answer-level preference direction, while Cloud scoring selects concrete trajectories and filters pairs by their score margin. Experiments on GSM8K and MATH-Numeric across four model settings show that Cloud-ScPO consistently improves over ScPO, with gains of up to 4.49% on GSM8K and 4.19% on MATH-Numeric. Pair-level analyses further show that Cloud-ScPO maintains comparable correctness reliability while more effectively separating informative chosen trajectories from incomplete, repetitive, or otherwise low-quality rejected responses.

Comments14 pages, 2 figures, 7 tables. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑