视觉语言模型何时应该观察?仅为需要且使用的视觉调用付费
When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used
浏览论文内容
中文总结 AI 辅助
针对视觉语言智能体视觉调用中大量虚假调用的问题,提出CounterCredit方法,通过决策价值和证据价值双重验证,仅奖励需要且使用的调用,显著提升基准性能并降低虚假调用率。
中文摘要 AI 辅助
视觉语言智能体在裁剪和缩放时,其训练奖励仅基于成功的工具调用,但一次成功的调用并不能证明模型需要观察或使用了所接收的像素。在我们的冷启动检查点上,仅10%至12%的视觉调用既是需要的又是使用的,而已发布的智能体在单个基准测试中有36%至87%的调用是虚假的。结果奖励、评判奖励和分支探针各自只观察到这一失败的一个侧面,并且结果奖励所支付的约三分之二用于既不需要也未使用的调用。CounterCredit在每次图像返回调用的实现前状态上提出这两个问题,使用策略自身的黄金答案分数。决策价值将实现的视觉分支与立即回答进行比较;证据价值将返回的裁剪与随机同尺寸补丁替换到同一调用中进行比较。通过两项验证的调用获得返现,而其他执行的调用需支付租金;价格有界,使得每个正确的轨迹都优于每个错误的轨迹,双通道GRPO优势将价格保持在其自身单位中。从相同的冷启动、提示池和预算出发,CounterCredit在V*上达到89.5%,在HR-Bench-4K上达到80.2%,在HR-Bench-8K上达到76.4%,比仅使用结果的GRPO高6.3至9.4个百分点,每问题调用次数为1.78对1.84,并将虚假调用率降至31%至36%,在所评估的智能体中最低。同样的方法将Qwen3-VL-8B基础模型从平均75.4提升至80.8。
英文摘要
Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it received. On our cold-start checkpoint only 10% to 12% of visual calls were both needed and used, and released agents make spurious calls 36% to 87% of the time on individual benchmarks. Outcome rewards, judge rewards, and branch probes each observe one side of this failure, and about two thirds of what an outcome reward pays goes to calls that were neither needed nor used. CounterCredit asks both questions of every image-returning call at its realized pre-call state, using the policy's own gold-answer score. A decision value compares the realized visual branch with answering immediately; an evidence value compares the returned crop with random same-size patches substituted into the same call. A call verified on both earns cashback and every other executed call pays rent; the price is bounded so that every correct trajectory outranks every wrong one, and a dual-channel GRPO advantage keeps the price in its own units. From the same cold start, prompt pool, and budget, CounterCredit reaches 89.5% on V*, 80.2% on HR-Bench-4K, and 76.4% on HR-Bench-8K, 6.3 to 9.4 points above outcome-only GRPO at 1.78 against 1.84 calls per question, and lowers the spurious-call rate to 31% to 36%, the lowest among the agents evaluated. The same recipe lifts a Qwen3-VL-8B base from 75.4 to 80.8 on average.
发表机构
- Tongji University(同济大学)
- Nankai University(南开大学)
- University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。