arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMPostTrainBench:多模态后训练自主研究的基准测试

MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training

Yuxin Liu, Yuxuan Wang, Zhenxin Lei, Lingchen Meng, Yuchong Sun, Junming Lin, Hongcheng Liu, Yunfei Chu, Qize Yang, Jin Xu, Lei Zhang, Zhendong Mao

arXiv 2610.05398首次发表:更新:

发表机构

University of Science and Technology of China; Alibaba Token Hub, Alibaba Group; University of the Chinese Academy of Sciences; Tsinghua University; Shanghai Jiao Tong University(中国科学技术大学; 阿里巴巴集团阿里云通义千问团队; 中国科学院大学; 清华大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MMPostTrainBench基准,评估LLM智能体在多模态后训练中的自主研究能力,发现性能回归与选择差距,并引入MMResearch框架以提升提交模型准确率。

AI 中文摘要

自主研究旨在通过迭代实验和反馈实现模型的持续改进。LLM智能体在自动化机器学习和语言模型后训练方面展现出潜力,但其维持多模态改进的能力仍不明确。我们提出了MMPostTrainBench,一个涵盖图像、音频、视频以及音视频联合理解与图像接地软件修复等八项任务的基准。智能体在固定预算内从共同的基础模型出发,利用开发反馈,然后对其提交的模型进行独立评估。评估涵盖目标与非目标模型结果、迭代模型改进与选择以及研究完整性。在所有八项任务中,52.1%的模型-任务均值低于基础模型,且评估的提交物也表现出非目标回归。模型性能在研究迭代中并未持续提升,智能体也未能可靠地选择最佳评估候选进行提交;最终提交物落后于该候选最多5.38个百分点。将自主研究从纯文本扩展到多模态任务,引入了感知、跨模态对齐和时间接地方面的额外错误源。观察到的回归和选择差距凸显了在目标改进与非目标能力保留之间取得平衡,并在研究迭代中保持收益的必要性。这些需求催生了MMResearch,一个多模态研究框架,它将媒体接地证据与假设和干预措施联系起来,通过分层记忆跨轮次携带发现,并利用开发评估保留候选。将其添加到现有代码智能体运行时后,对于Claude Opus 4.8与Claude Code,提交模型的准确率最多提升7.75个百分点;对于GPT-5.6-sol与Codex,提升2.33个百分点。

英文摘要

Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.

Comments17 pages. Code: https://github.com/sod1010/MMPostTrainBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑