发表机构
Zhejiang University; Ant Group; Hong Kong Baptist University; Nanyang Technological University; Southeast University(浙江大学; 蚂蚁集团; 香港浸会大学; 南洋理工大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DE-Venus 是一种高效数据 RLVR 统一框架,通过三个模块实现监督管理,仅用 10% 标签或 13% 数据即可维持/提升模型质量,减少收敛步数,降低标注与训练成本。
AI 中文摘要
带可验证奖励的强化学习(RLVR)可提升大语言模型的推理能力,但其实际规模化应用受限于昂贵的在线策略 rollout 及大规模获取可靠目标的成本。现有方法分别处理样本选择、不完整监督或噪声标签,常将监督逻辑与分布式训练绑定,阻碍了可控比较与复用。本文提出 DE-Venus,一种面向高效数据 RLVR 的统一框架,将监督视为数据准备与策略优化过程中演化的状态,将该生命周期划分为三个模块:主动数据选择分配训练与标注预算;弱监督构建从未标注样本中推导学习信号;训练时监督优化过滤或修正不可靠的监督。DE-Venus 支持七种代表性方法及数据选择流水线,通过将特定方法的决策表示为离线数据集转换或目标、奖励、批次与优势的在线转换,同时保留 verl 的分布式执行约定。在公开基准及三个业务场景中,独立配置仅使用 10% 的标签或低至 13% 的相关数据即可维持或提升模型质量;选定的业务配置还将观测到的收敛步数减少了 63%--75%。因此,DE-Venus 在不牺牲可扩展 RL 执行的前提下降低了标注与训练成本。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, often entangling supervision logic with distributed training and hindering controlled comparison and reuse. We present DE-Venus, a unified framework for data-efficient RLVR that treats supervision as evolving state across data preparation and policy optimization. It organizes this lifecycle into three modules: Active Data Selection allocates training and annotation budgets; Weak Supervision Construction derives learning signals from unlabeled examples; and Training-Time Supervision Refinement filters or corrects unreliable supervision. DE-Venus supports seven representative methods and a data-selection pipeline by expressing method-specific decisions as offline dataset transitions or online transformations of targets, rewards, batches, and advantages while preserving verl's distributed execution contracts. Across public benchmarks and three business scenarios, separate configurations preserve or improve model quality with only 10% of labels or as little as 13% of relevant data; selected business configurations also reduce observed convergence steps by 63%--75%. DE-Venus thus reduces annotation and training costs without sacrificing scalable RL execution.