arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RISED:面向智能体多环境选择与自蒸馏的评分准则

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu

arXiv 2610.00979首次发表:更新:

发表机构

National University of Singapore; Apple(新加坡国立大学; 苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RISED利用跨环境共享的评分准则标签,指导多环境智能体强化学习中的在线数据选择与自蒸馏监督,提升各环境平均通过率。

AI 中文摘要

跨多种交互环境联合训练单个LLM智能体作为通向通用智能体的途径,已日益受到关注。现有的课程学习与数据选择策略往往在环境层面分配训练或优先考虑基于局部奖励的信号,而未明确考虑不同环境间当前轨迹的关系以进行提示组选择。同时,由于各环境的学习速率不同,一个批次内可能同时存在全部失败与全部成功的轨迹组,导致这些数据缺乏组内相对奖励信号。上述挑战凸显了在多环境强化学习中仅依赖标量奖励的局限性:标量奖励在跨环境关系方面提供的信息有限,且在奖励相同时缺乏组内奖励对比。这促使我们利用更丰富的文本反馈(如描述轨迹行为的评分准则)来指导学习。除了将评分准则用作奖励外,我们重新利用评分准则来同时指导在线数据选择与策略监督。一个LLM评判器使用跨环境共享的预定义评分准则词汇表为每条轨迹打标签。由此产生的画像用于指导选择与混合环境批次整体行为构成相符的数据,同时限制与已选数据的重叠。可用的正面评分准则(描述期望行为)为在线策略自蒸馏教师提供特权上下文,补充了词元级别的监督,而负面评分准则(描述非期望行为)则引导后续轨迹生成远离反复出现的失败模式。这些组件共同构成了RISED。在不同模型骨干上,RISED在各类环境中取得了最高的平均通过率,并在每个单独环境中排名第一或第二。基于评分准则的RISED分析可进一步刻画伴随这些提升的行为变化。

英文摘要

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics' usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑