arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09657cs.CVcs.AIcs.MM

用于语言智能的可扩展视觉预训练

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Huanze Tang, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haij… 展开作者

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Huanze Tang, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究挑战语言模型仅在文本表示上训练的假设,通过对直接利用视觉文档的无监督视觉预训练范式进行系统研究,发现其在多主干和基准上优于仅文本预训练,为可扩展语言智能提供有效途径。

中文摘要 AI 辅助

大型基础模型的快速发展主要得益于对大规模文本语料库的预训练。然而,许多知识通过视觉表示传达,图形、排版方程和页面布局携带的丰富信息无法仅靠文本完全捕捉。当前预训练方法通过将文档和网页等视觉丰富的源转换为纯文本学习语言智能,从而丢弃了这些视觉线索。本文挑战语言模型必须仅在文本表示上训练的默认假设,表明视觉预训练是基础模型智能的可扩展学习者。为此,我们对直接利用视觉文档而不进行文本提取的无监督视觉预训练范式进行了系统研究。在多个主干和基准上,对相同基础语料库的视觉预训练始终优于仅文本预训练,为可扩展语言智能提供了有效途径。

英文摘要

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.

↑