arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像素文本表示学习的设计基础

On the Design Fundamentals of Pixel Text Representation Learning

Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao

arXiv 2609.01147首次发表:更新:

发表机构

The Chinese University of Hong Kong; DAMO Academy, Alibaba Group; Fudan University; Aberystwyth University; University of Macau; Shanghai University of Finance and Economics(香港中文大学; 阿里巴巴达摩院; 复旦大学; 阿伯里斯特威斯大学; 澳门大学; 上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对现有像素-文本编码器的缺陷,确定视觉文本表示学习的四个关键设计原则,据此训练的Pixel Linguist II在多项视觉文本任务上达SOTA,且在高视觉令牌压缩下仍具鲁棒性,推动光学上下文压缩发展。

AI 中文摘要

富含文本的视觉输入需要能直接在像素空间中读取、检索和压缩语言的模型,但现有的像素-文本编码器存在固定分辨率预训练、视觉捷径学习、视觉定位能力弱、多语言视觉文本理解不足等问题。本研究探讨鲁棒视觉文本表示学习所需的基本设计原则,通过系统受控消融实验确定四个关键组件:可变图像分辨率和渲染字体大小为高分辨率文档泛化提供空间代理;自然图像-文本对对定位不可或缺,可避免仅文本崩溃;布局感知渲染有助于防止像素级捷径;两阶段多语言课程实现有效的跨语言对齐。将这些原则整合为可扩展训练方案后,我们训练了Pixel Linguist II,这是一种原生分辨率视觉编码器,采用即时渲染、统一对比定位及2.8亿训练样本的多语言课程进行训练。Pixel Linguist II在英语、跨语言和多语言视觉语义文本相似度(Visual STS)及ViDoRe任务上达到新的SOTA,还能在多模态大语言模型(MLLM)下游评估中表现更优;值得注意的是,Pixel Linguist II在80%视觉令牌压缩下仍保持鲁棒性,在光学上下文压缩领域展现出巨大潜力。我们的代码和资源可在该https URL获取。

英文摘要

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual text understanding. In this work, we investigate the fundamental design principles required for robust visual text representation learning. Through systematic controlled ablations, we identify four critical components: variable image resolutions and rendered font sizes provide spatial proxies for high-resolution document generalization; natural image-text pairs are indispensable for grounding and prevent text-only collapse; layout-aware rendering helps prevent pixel-level shortcuts; and a two-stage multilingual curriculum enables effective cross-lingual alignment. By integrating these principles into a scalable training recipe, we train Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples. Pixel Linguist II sets new state-of-the-art results on English, cross-lingual, and multilingual Visual STS and ViDoRe, while also enabling better MLLM downstream evaluation. Notably, Pixel Linguist II remains robust under 80\% visual token compression, showing great promise for optical context compression. Our code and resources are available at https://github.com/Pixel-Linguist/Pixel-Linguist-II.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑