arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估条件化训练:让模型泛化到更强的监督机制

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

Alec Harris, Kasey Corra, Archie Chaudhury, Yixiong Hao

arXiv 2608.10209首次发表:更新:

发表机构

AI Safety Initiative at Georgia Tech; University of Chicago(佐治亚理工学院AI安全倡议; 芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ECT后训练框架,作为SFT、PPO的附加组件,通过关联反馈保真度提升不完美反馈下的模型性能,在新闻生成和算术任务中验证其可改善目标行为。

AI 中文摘要

用于训练大型语言模型(LLMs)的反馈信号是决定其行为的主要因素,也是我们使其与人类价值观和目标对齐的主要手段。然而,当前后训练方法的一个关键局限在于,人类标注者和自动奖励函数无法忠实地捕捉到我们想要给出的反馈。我们提出评估条件化训练(Evaluation-Conditioned Training,ECT),这是一种后训练框架,它使用自然语言让每个训练样本与我们提供的反馈保真度相关联,然后在部署时通过让LLM依赖高保真监控器来引出期望的行为。ECT旨在改进在不完美反馈下的性能,可作为SFT、PPO等现有算法的附加组件使用。我们首先为ECT提供概念框架,并讨论其解决奖励错误指定持续来源的潜力;接着在引出潜在知识(eliciting latent knowledge,ELK)问题的背景下说明ECT的动机;最后在两个概念验证实验中对ECT进行评估:一是提高新闻文章生成的公正性,二是减少算术任务中的奉承行为。在每个设置中,我们分别使用不完美反馈(奖励偏见)和不完美反馈(奖励与用户的一致性),在这两种设置中,ECT相较于直接训练都改进了目标行为。

英文摘要

Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post-training methods is the inability of human annotators and automated reward functions to faithfully capture the feedback we would like to give. We introduce Evaluation-Conditioned Training (ECT), a post-training framework that uses natural language to condition each training sample on the fidelity of the feedback we provide and then elicits the desired behavior by conditioning the LLM on a high-fidelity monitor in deployment. ECT is aimed at improving performance under imperfect feedback and works as an add-on to existing algorithms such as SFT and PPO. We first provide a conceptual framework for ECT and discuss its potential to address persistent sources of reward mis-specification. Then we motivate ECT in the context of the eliciting latent knowledge (ELK) problem. Finally, we evaluate ECT on two proof-of-concept experiments: increasing even-handedness in news article generation and reducing sycophancy on an arithmetic task. In each setting, we utilize imperfect feedback, rewarding bias and agreement with the user, respectively. In both settings, ECT improves the targeted behavior relative to direct training.

Comments16 pages, 7 figures. Accepted at the Agent Behavior Workshop at COLM 2026. Code: https://github.com/evaluationconditionedtraining/Evaluation-Conditioned-Training

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑