发表机构
University of Rochester; Univeristy of Michigan(罗切斯特大学; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究主动视觉感知智能体在不同预算下的任务执行问题,提出AdaTurn框架,通过强制答案DAPO等组件,根据预算调整智能体行为,提高低预算准确率,且能有效转移到多基准测试中。
AI 中文摘要
主动视觉智能体通过在多轮中交错推理与图像定位动作来解决细粒度图像任务。然而,部署时的展开预算很少是固定的。现有方法在训练策略时未考虑预算,导致灾难性截断。本文提出AdaTurn,一个预算感知框架,根据允许的轮数调整智能体,并明确训练由预算引起的边界行为。关键组件强制答案DAPO将超预算事件转化为可训练的最终决策步骤。训练和推理时随机化展开预算,并引入负载平衡调度器。AdaTurn显著提高了低预算准确率,如在四回合时将VisualProbe-Medium从36.7%提高到47.6%,同时在更大预算下保持良好缩放,并有效转移到多个主干和通用多模态基准测试中。
英文摘要
Active visual agents solve fine-grained image tasks by interleaving reasoning with image-grounding actions across multiple turns. However, deployment-time rollout budgets are rarely fixed: some requests permit long rollouts, while others require the agent to act under a tight turn limit. Existing methods train the policy as if the rollout budget were hidden, so when the available budget is smaller than the trajectory the agent prefers, the interaction is often truncated before any valid answer is produced; we term this failure \emph{catastrophic truncation}. To overcome this challenge, we present AdaTurn, a budget-aware framework that conditions the agent on the allowed number of turns and explicitly trains the boundary behavior induced by the budget. Our key component, Forced-Answer DAPO (FA-DAPO), converts the over-budget event from a masked or penalized failure into a trainable final-decision step, teaching the model to synthesize partial evidence when further tool use is no longer possible. We further randomize rollout budgets during both training and inference and introduce a load-balanced scheduler that makes such operations practical. AdaTurn substantially improves low-budget accuracy, for example raising VisualProbe-Medium from 36.7% to 47.6% at four turns, while preserving strong scaling at larger budgets and transferring effectively to multiple backbones and general multimodal benchmarks.