arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26873cs.CL

SERPO:面向开放式测试时强化学习的自进化规则策略优化

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Jianze Wang, Kunwang Zheng, Ying Liu, Yu Cao, Qilong Zhang, Jinlong Chen, Hua Yang, Qianglong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

SERPO通过协同进化响应证据、查询规则和策略参数的闭环机制,在开放式测试时强化学习任务中显著提升了HealthBench等基准的性能,支持分布外迁移与跨基准进化。

中文摘要 AI 辅助

测试时强化学习(TTRL)使语言模型能在推理时无需标注反馈实现自进化。现有方法依赖答案投票,因此无法自然扩展至开放式生成——这类任务中有效响应无法映射到统一标准答案。在无外部奖励模型或更强评判者的情况下,自适应过程必须从模型自身输出中构建可靠奖励信号。我们提出SERPO(Self-Evolving Rubric Policy Optimization,自进化规则策略优化),它用一个闭环取代答案投票,该闭环协同进化响应证据、查询特定规则和策略参数。Good-Normal-Bad(G-N-B,好-中-差)响应进化将最大分离的rollout组织为有序档案;规则进化保留能区分这些档案的标准;概率标准评分将 verdict-token( verdict 令牌)似然转化为奖励信号;策略进化用所得信号优化Actor。新的Actor rollout随后刷新档案和规则,形成三方进化闭环。在两种模型配置、两个域内基准和四个OOD(分布外)基准上,SERPO使HealthBench和ResearchQA较对应基础模型分别提升最高20.63和20.31个点,使六个基准的宏平均提升最高8.06个点,且支持分布外迁移和跨基准持续进化。

英文摘要

Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.

发表机构

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑