AI 中文总结
针对长程多模态推理中策略与技能不匹配的问题,提出RLHARNESS框架,通过版本化Harness交替演化技能与策略,结合重建与DAPO II,显著提升准确率和F1分数。
AI 中文摘要
多模态推理要求模型在长决策链中保留视觉证据,同时在不同场景和规则下选择合适的程序。当学习仅由终端验证器引导时,强化学习(RL)能揭示最终答案是否正确,但无法说明应如何生成答案。因此,策略必须在学习执行可复用推理程序的同时发现这些程序,从而产生程序冷启动问题。技能可以将成功的程序外部化,减少重复探索,并提供可检查的指导。然而,固定的技能库假设该指导与不断演化的策略保持兼容,而仅更新技能可能导致其触发条件、执行协议和演示样例过时或相互不一致。我们提出RLHARNESS,将技能、选择与执行协议、少样本演示和任务契约组织成一个统一的、带版本管理的Harness,并交替进行Harness演化与策略学习。探索-蒸馏Harness构建初始Harness以及用于SFT和DAPO I的版本对齐的验证轨迹。在第一个RL块之后,后RL重建Harness从新的成功-失败轨迹中重建技能、协议和演示,DAPO II则使策略适应重建后的程序。RLHARNESS在MetroMap/TravelMap上将准确率从16.25%/27.50%提升至62.00%/50.00%,在Fee-VL/Cancel-VL上将F1分数从37.13%/45.50%提升至65.81%/65.51%。所有四个任务仅在重建和DAPO II之后取得最佳结果,表明演化的Harness通过持续更新策略学习执行的外部程序来补充强化学习。
英文摘要
Multimodal reasoning requires models to preserve visual evidence through long decision chains while selecting appropriate procedures across diverse scenarios and rules. When learning is guided only by terminal verifiers, reinforcement learning (RL) reveals whether a final answer is correct but not how it should be produced. The policy must therefore discover reusable reasoning procedures while learning to execute them, creating a program cold-start problem. Skills can externalize successful procedures, reduce repeated exploration, and provide inspectable guidance. However, a fixed Skill Bank assumes that this guidance remains compatible with an evolving policy, while updating Skills alone can leave their triggers, execution protocols, and demonstrations stale or mutually inconsistent. We introduce RLHARNESS, which organizes Skills, selection and execution protocols, few-shot demonstrations, and task contracts into a unified, versioned Harness and alternates Harness evolution with policy learning. An Exploration-Distillation Harness builds the initial Harness and version-aligned verified traces for SFT and DAPO I. After the first RL block, a Post-RL Reconstruction Harness rebuilds Skills, protocols, and demonstrations from fresh success-failure rollouts, and DAPO II adapts the policy to the reconstructed program. RLHARNESS improves Accuracy from 16.25%/27.50% to 62.00%/50.00% on MetroMap/TravelMap and raises F1 score from 37.13%/45.50% to 65.81%/65.51% on Fee-VL/Cancel-VL. All four tasks achieve their best results only after reconstruction and DAPO II, showing that an evolving Harness complements RL by continually updating the external program that the policy learns to execute.