通过对比课程学习实现高效缩放的方法
Learning to Zoom Efficiently with a Contrastive Curriculum
浏览论文内容
中文总结 AI 辅助
本文针对视觉智能体的缩放任务,提出无需监督微调的对比课程内在奖励方法,在多个基准测试中表现更优,还引入M&C数据集用于评估缩放能力。
中文摘要 AI 辅助
使用缩放工具是现代视觉智能体的重要基础组成部分,因为它能高效处理涉及高分辨率图像的任务。此前多数方法需要大量的预启动监督微调(SFT)阶段来教导模型进行缩放,本文证明这并非必要,提出了一种用于多模态大语言模型(MLLM)中学习工具使用的新型内在奖励,无需额外标签或预启动SFT。该类InfoNCE的奖励采用难度逐步提升的负工具调用课程作为对比训练信号。在$V^*$、HRBench和MME-RealWorld上的实证实验表明,本文方法具有竞争力且效率更高;当作为SFT的替代方案使用时,其性能甚至优于所有基线。为直接测量模型的缩放能力,本文进一步引入了可扩展的合成Muffin&Chihuahua(M&C)数据集,每个图像由网格组成,每个单元格显示松饼或吉娃娃。利用M&C数据集独特的感兴趣区域标签,本文发现召回率是与缩放区域和最终任务性能相关性最强的指标。本文的模型和复现代码可在该httpsURL公开获取。
英文摘要
Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images. Most previous methods need an extensive warm-start supervised fine-tuning phase for teaching models zoom-in. We show that this is not necessary by proposing a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT. Our InfoNCE-style reward uses a curriculum of increasingly hard negative tool calls as a contrastive training signal. Empirical experiments on $V^*$, HRBench and MME-RealWorld show that our approach is competitive while being more efficient. When used as a drop-in replacement for SFT, we even outperform all baselines. To directly measure the zoom-in ability of models, we further introduce the scalable synthetic Muffin&Chihuahua (M&C) dataset. Each image consists of a grid with every cell either showing a muffin or chihuahua. Leveraging the M&C dataset's unique region of interest labels, we find that recall is the metric that most strongly correlates the zoom-in region with final task performance. Our model and code for reproduction is publicly available under https://github.com/UKPLab/emnlp2026-zoom-in
发表机构
- Technical University of Darmstadt(达姆施塔特工业大学)
- Hessian Center for AI(黑森州人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。