Trinity:具有自验证器的自进化视觉语言模型
Trinity: Self-Evolving Vision-Language Models with a Self-Verifier
浏览论文内容
中文总结 AI 辅助
提出Trinity框架,让一个VLM同时扮演提问器、求解器和自验证器,通过自验证器筛选问题和答案,仅用未标注图像训练,显著提升科学推理能力。
中文摘要 AI 辅助
自进化视觉语言模型(VLM)是一种自我改进的形式,其中模型从未标注的图像中生成自己的训练数据,这是一种有前景的途径,使智能体能够以无监督的方式扩展其推理能力,而无需依赖不断增加的标注预算。现有方法将一个提出问题的提问器(Questioner)与一个回答问题的求解器(Solver)配对,但主要通过采样答案之间的一致性来奖励这两个角色。一致性是真理的弱代理:它无法判断问题是否基于图像,无法判断提议的参考答案是否正确,也无法判断自信的多数是否以相同方式出错。我们提出了Trinity,其中一个VLM扮演三个角色:提问器、求解器和验证器(Verifier),其中验证器是一个自验证器:策略本身的指数移动平均(EMA),既不需要标签也不需要外部评判者。验证器在每道生成的问题成为监督信号之前,筛选其图像基础性和答案正确性,根据图像对求解器的推理进行评分,并裁决参考答案与强大的求解器共识之间的争议,在共识正确时纠正参考答案并惩罚提问器。仅使用图像进行训练,Trinity在数学视觉推理和包含生物学内容的科学基准上提升了Qwen3-VL-8B的性能,例如,在SciVQR的生物学子集上提高了+8.6分,在MathVerse上提高了+12.8分,其奖励动态表现如同健康的自我对弈课程。这些结果表明,一个自进化的多模态智能体可以仅从未标注的科学图像中增强其科学推理能力,而模型本身充当验证器。
英文摘要
Self-evolving vision-language models (VLMs), a form of self-improvement in which a model generates its own training data from unlabeled images, are a promising route toward agents that expand their reasoning capability in an unsupervised manner, without relying on ever-larger annotation budgets. Existing methods pair a Questioner that proposes problems with a Solver that answers them, but reward both roles mainly by agreement among sampled answers. Agreement is a weak proxy for truth: it cannot tell whether a question is grounded in the image, whether the proposed reference answer is right, or whether a confident majority is wrong in the same way. We present Trinity, in which one VLM plays three roles, Questioner, Solver, and Verifier, and the Verifier is a self-verifier: an exponential moving average (EMA) of the policy itself, requiring neither labels nor an external judge. The Verifier screens every generated question for image grounding and answer correctness before it becomes supervision, scores Solver reasoning against the image, and adjudicates disputes between the reference answer and a strong Solver consensus, correcting the reference and penalizing the Questioner when the consensus is right. Trained on images alone, Trinity improves Qwen3-VL-8B on mathematical visual reasoning and on science benchmarks with biology content, for example, +8.6 points on the biology split of SciVQR and +12.8 on MathVerse, and its reward dynamics behave as a healthy self-play curriculum should. These results suggest that a self-evolving multimodal agent can strengthen its scientific reasoning from unlabeled scientific images alone, with the model itself serving as the verifier.
发表机构
- Electronics and Telecommunications Research Institute (ETRI)(韩国电子通信研究院)
- Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。