arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一多模态模型中的递归自我改进

Recursive Self-Improvement in Unified Multimodal Models

Huijuan Wang, Chufan Shi, Cheng Yang, Yaokang Wu, Taylor Berg-Kirkpatrick, Xuezhe Ma

arXiv 2610.03002首次发表:更新:

发表机构

University of Southern California; University of California San Diego; Carnegie Mellon University(南加州大学; 加利福尼亚大学圣迭戈分校; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出递归跨能力自我改进(RSI)方法,使统一多模态模型通过程序执行作为外部真值,相互训练文本与视觉能力,在图表基准上将得分从45.7%提升至60.2%,显著优于持续训练。

AI 中文摘要

统一多模态模型(UMMs)能够理解和生成文本与图像,这使得模型能够生成自己的训练数据。现有的UMMs自我改进方法在视觉侧保持监督,即通过图像理解来评判图像生成。我们提出了递归跨能力自我改进(RSI),这是一种训练循环,其中UMM的文本和视觉能力相互提供训练数据。在每一轮中,模型生成图像并读取它们以发现自身的不足。然后,它针对这些不足编写程序,并通过执行来验证每个结果是否符合其规格。经过验证的渲染结果用于训练图像生成,而标注的渲染结果和模型自身的正确程序则用于训练视觉理解和程序编写。因此,程序执行充当了模型外部的真值来源,使得错误不会在轮次间累积。我们在图表上研究了RSI,并构建了BasicChartBench来评估训练早期的开放模型。在与训练不同的措辞请求下,四轮RSI将得分从45.7%提升至60.2%,而持续训练则保持在46.3%。经过验证的构建过程贡献了大部分增益,而针对模型失败进行优化额外增加了3.5%。在此过程中,经过验证的程序比例从48.9%上升至95.2%,读者在编辑后的渲染结果上的准确率从55.6%提升至87.4%。

英文摘要

Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑