用于视频推理的多智能体自增强强化学习
Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning
浏览论文内容
中文总结 AI 辅助
该研究提出结合可训练Grounder与冻结Verifier的多智能体强化学习框架,在20亿参数规模下实现视频推理任务零样本迁移,相关指标优于同等规模基线,验证了冻结验证作为证据选择训练信号的有效性。
中文摘要 AI 辅助
视频推理任务(如基于视频的问答和时序定位)需要选择支持查询的时序证据。在许多当前的训练设置中,时序监督通过局部目标(如边界回归或片段生成)实现,而验证主要用于推理时对候选片段进行重新排序。本文研究冻结的验证器是否也能指导训练。我们的多智能体框架将可训练的Grounder(定位器)与冻结的Verifier(验证器)结合:Grounder采样候选轨迹和证据片段,Verifier分配查询条件下的片段分数;组相对策略梯度目标有利于优于输入内其他轨迹的样本,自举校准损失引导时序预测朝向验证器偏好的片段。在源任务上训练且未对目标数据集微调的情况下,一个20亿参数的实例在基于视频的问答、时序定位和长视频问答上实现零样本迁移,在基于视频的问答基准上达到28.7%的交并比和25.4%的答案定位准确率,在时序定位基准上达到46.1%的交并比,在长视频问答基准上达到54.1%的准确率。与同等规模的强基线相比,增益虽小但一致,在交并比和中等重叠召回等相关性导向指标上提升最明显。在测试的基准和迁移设置内,结果支持冻结验证作为证据选择的训练信号,同时表明严格边界精度仍相对较弱。代码和模型可在该httpsURL获取。
英文摘要
Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform their within-input peers, and a bootstrapped calibration loss steers temporal predictions toward verifier-preferred spans. Trained on source tasks and evaluated without target-dataset fine-tuning, a two-billion-parameter instantiation transfers zero-shot across grounded question answering, temporal grounding, and long-video question answering, reaching 28.7\% intersection-over-union and 25.4\% answer-grounding accuracy on a grounded-question-answering benchmark, 46.1\% intersection-over-union on a temporal-grounding benchmark, and 54.1\% on a long-video question-answering benchmark. Relative to a strong same-scale baseline, the gains are modest but consistent, with the clearest improvements on relevance-oriented metrics such as intersection-over-union and moderate-overlap recall. Within the tested benchmarks and transfer setting, the results support frozen verification as a training signal for evidence selection, while showing that strict boundary precision remains comparatively weaker. Code and models are available at https://anonymous.4open.science/r/MASIRL-E50C/
发表机构
- Lanzhou University(兰州大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。