发表机构
Duke University; Netflix(杜克大学; 网飞公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对主观任务中LLM推理的失败模式,提出带条件长度惩罚的后训练算法缓解推理崩溃,还发现推理风格不匹配是主观验证误差主因,给出算法补丁与架构蓝图。
AI 中文摘要
推荐系统依赖个性化,其中“正确性”极少是二元事实,而是人类主观偏好的问题。当大语言模型(LLM)被部署为安全与质量准则的自主验证器时,它们面临一个独特挑战:感知上下文的偏好对齐。近期可验证奖励强化学习(RLVR)的进展大多针对客观数学任务。通过在生产推荐平台上针对四个真实世界验证任务,对专有与开源模型开展的大规模研究,我们探究显式推理是否能泛化到以人类为中心的主观行业准则。我们揭示了一个根本漏洞:僵化的、以数学为中心的推理轨迹会显著降低验证性能,而应用标准RLVR会触发一种我们称为“推理崩溃”的现象,即策略放弃深思熟虑,转而采用快速启发式猜测。我们引入一种带条件长度惩罚的后训练算法,该算法将验证准确率与受限推理长度相结合,可终止崩溃并恢复性能。最后,我们表明推理轨迹的功效与其社会语言学框架紧密相关:在1500个合成角色中,仅因采用的推理角色不同,验证准确率的宏F1值波动近0.38——这表明许多主观验证误差实际是推理风格不匹配。这一观察催生了一种中间训练架构,该架构通过上下文对齐的角色来路由推理。本工作既提供了可扩展的算法补丁,也提供了使推理模型与现实世界主观约束对齐的长期架构蓝图。
英文摘要
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality guidelines, they face a distinctive challenge: context-aware preference alignment. Recent gains in Reinforcement Learning with Verifiable Rewards (RLVR) are indexed mostly on objective, mathematical tasks. Through a large-scale study spanning both proprietary and open-source models on four real-world verification tasks from a production recommender platform, we ask whether explicit reasoning generalizes to subjective, human-centric industry rubrics. We expose a fundamental vulnerability: rigid, math-centric reasoning traces actively degrade verification, and applying standard RLVR triggers a phenomenon we term reasoning collapse, in which the policy abandons deliberation in favor of rapid heuristic guessing. We introduce a conditional length-penalized post-training algorithm that intertwines verification accuracy with bounded reasoning length, halting collapse and recovering performance. Finally, we show that a reasoning trace's efficacy is tightly coupled with its socio-linguistic framing: across 1500 synthesized personas, verification accuracy swings by nearly 0.38 macro-F1 depending solely on the adopted reasoning persona---evidence that much subjective-verification error is really reasoning-style mismatch. This observation motivates a mid-training architecture that routes reasoning through contextually aligned personas. This work offers both a scalable algorithmic patch and a long-term architectural blueprint for aligning reasoning models with real-world subjective constraints.
Comments20th ACM Conference on Recommender Systems (RecSys 2026)