AI 中文总结
本文通过STANCE-BENCH基准识别测试时缩放和训练在个体立场预测中的四种失败模式,并提出结合直接得分与历史支持评估的方法,在Qwen3-8B上提升宏F1至21.83。
AI 中文摘要
测试时缩放和训练后处理已提升了大型语言模型在编码和数学推理方面的性能,但它们对个体立场预测的有效性仍不明确。我们通过从一个人的历史记录中预测其在新讨论中的立场来研究这一问题。我们评估了广泛使用的测试时缩放策略和训练后处理方法,如监督微调和强化学习,并识别出生成、选择和训练中的四种失败模式:(1)错误共识,即重复采样一致地得出错误立场;(2)选择失败,即生成覆盖了观察到的立场但选择却遗漏了它;(3)响应过拟合,即监督微调改善了模仿但损害了预测;(4)早期平台期,即强化学习表现出适度的初始提升但后续改进有限。我们使用STANCE-BENCH揭示了这些失败,该基准包含来自500名Hacker News用户的2499个预测任务。在此分析的指导下,我们探索了一种简单方法,该方法结合了所有候选立场的直接得分以及对个体历史支持度的明确评估。在781个任务的测试集上,该方法使用Qwen3-8B实现了21.83的讨论特定宏F1分数,而直接得分为19.27。我们的结果促使分别评估候选生成、最终选择和个人特定证据的使用。
英文摘要
Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately. Our data is available at https://github.com/stance-bench/Stance-Bench.