arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32196cs.CLcs.LG

裁判并非其孪生:后训练使模型的写作更具可预测性,但作为裁判,其品味几乎未向可预测写作倾斜

The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writing

Arman Nik Khah, Arvin Bahreini

首次发表
浏览论文内容

中文总结 AI 辅助

本研究追踪OLMo-2和Zephyr模型的后训练阶段,发现作为作者时写作可预测性显著增加,但作为裁判时对可预测写作的偏好几乎不变,表明自动化评估仍可感知创造力进步。

中文摘要 AI 辅助

语言模型现在通常由其他语言模型进行评分。如果后训练使模型自身的写作更具可预测性,那么同一模型作为裁判时也可能学会奖励可预测的写作,从而使得创造力方面的进步在自动化评估中不可见。我们跟踪了两个开放模型系列OLMo-2和Zephyr(各7B参数),贯穿其公开的训练阶段,并在每个阶段进行两次测量:一次作为短篇故事的作者,一次作为故事对的裁判。作为作者,模型的漂移正如所担忧的那样:每个系列完全训练后的模型发现其基础、监督微调(SFT)和偏好训练(DPO)阶段的故事在所有十个提示下都越来越熟悉,而来自另一个系列的模型发现训练后的故事每个词元的意外程度降低了7.0%至9.2%(OLMo-2)和24%至25%(Zephyr)。作为裁判,它们几乎未向可预测写作倾斜。当被问及哪个故事更好时——这是每个裁判都可以用来区分故事与其乱序词的问题——没有训练后的裁判对更可预测故事的估计倾向增长达到一个百分点(在选择它的概率上)。事后单侧95%上限显示,在平均故事对上该增长为2.7个百分点,大约相当于未训练的OLMo-2裁判自身的倾向。当被问及哪个更具创造力时,没有训练后的裁判的估计比其基础模型更偏向可预测故事。训练反而强化了在创造力问题下对更长故事的偏好,以及对一个答案槽位的偏好,并且破坏了“更具创造力”这一问题:被问及此问题的训练后裁判不再可靠地偏好一个故事胜过其乱序词。后续实验无法构建在可预测性上不同但在质量上相同的配对,因为使该作者的故事更不可预测的路径也破坏了其中一些故事,其频率足以使预先设定的质量下限失败。

英文摘要

Language models are now routinely graded by other language models. If post-training makes a model's own writing more predictable, it may also teach the same model, acting as a judge, to reward predictable writing, so that progress on creativity would be invisible to automated evaluation. We follow two open model families, OLMo-2 and Zephyr (7B parameters each), through their public training stages and measure every stage twice, as a writer of short stories and as a judge of pairs of stories. As writers, the models drift as feared: each family's fully trained model finds the stories of its base, supervised fine-tuned (SFT) and preference-trained (DPO) stages progressively more familiar, in all ten prompts, and a model from the other family finds the trained stories 7.0 to 9.2 percent (OLMo-2) and 24 to 25 percent (Zephyr) less surprising per token. As judges, they barely move toward predictable writing. Asked which story is better, a question every judge can use to tell a story from its own words in scrambled order, no trained judge's estimated tilt toward the more predictable story grows by as much as one point in the probability of picking it. A post hoc one-sided 95% upper bound on that growth is 2.7 points on an average pair, about the size of the untrained OLMo-2 judge's own tilt. Asked which is more creative, no trained judge's estimate favors the predictable story more than its base's does. Training instead strengthens a preference for longer stories when the question is creativity, and for one answer slot, and it breaks "more creative" as a question: trained judges asked it no longer reliably prefer a story to its scrambled words. A follow-up could not build pairs that differ in predictability but not in quality, because the routes that made this writer's stories less predictable also broke some of them, often enough to fail a quality floor set in advance.

补充信息

↑