AdaptEvo:结合演化监督的自适应智能体学习
AdaptEvo: Adaptive Agent Learning with Evolving Supervision
AI总结:
针对规则约束决策任务的不完善监督问题,提出AdaptEvo框架,结合CA-GRPO与演化决策知识,在工业多模态内容审核数据集上,较GRPO在同期、跨期测试集均实现性能提升。
AI中文摘要:
受规则约束的上下文决策任务要求模型将指定规则应用于特定案例的上下文与证据。书面规则可能在决策指导和过程评估中存在缺口,而参考判断对规则和证据的支持程度各不相同。为应对这些挑战,我们提出AdaptEvo,这是一种在不完善监督下进行学习的框架,它将置信度自适应策略优化与演化的决策知识及评估规则相结合。其训练模块采用置信度自适应GRPO(CA-GRPO),根据参考置信度平衡结果奖励与过程奖励;演化模块则从训练案例的反复失败中综合可复用的决策知识,并优化过程规则以检测被忽视的错误。为支持实证评估,我们构建了一个工业多模态内容审核数据集,包含训练集、同期测试集和跨期测试集,其中跨期测试集是在规则变更的情况下收集的。使用Qwen3.6-35B-A3B模型,AdaptEvo在同期测试集上达到61.9%的精确标签准确率和72.2%的二分类决策准确率,分别比GRPO高出7.5和3.7个百分点;在跨期测试集上,经CA-GRPO训练的策略在所有评估检查点中,在未注入决策知识的情况下,相较于基础模型保持了精确标签准确率的优势,而GRPO则随训练持续出现性能下降,且CA-GRPO在跨期测试集的两项指标上均优于所测试的固定奖励混合方案。
英文摘要:
Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence. Written rules can leave gaps in decision guidance and process evaluation, while reference judgments vary in their support from the rules and evidence. To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics. Its Training module uses Confidence-Adaptive GRPO (CA-GRPO) to balance outcome and process rewards according to reference confidence. Its Evolution module synthesizes reusable decision knowledge from recurring failures across training cases and refines process rubrics to detect overlooked errors. To support empirical evaluation, we construct an industrial multimodal content moderation dataset comprising a training set and In-Period and Out-of-Period test sets, with the latter collected under changed rules. Using Qwen3.6-35B-A3B, AdaptEvo achieves 61.9% exact-label accuracy and 72.2% binary decision accuracy on In-Period, exceeding GRPO by 7.5 and 3.7 percentage points, respectively. On Out-of-Period, the policy trained with CA-GRPO retains exact-label accuracy gains over the base model across evaluated checkpoints without injected decision knowledge, while GRPO declines with continued training. CA-GRPO also outperforms the tested fixed reward mixtures on both Out-of-Period metrics.