arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同语义,不同路径:面向视觉-文本压缩的自改进对齐

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

Tianyu Liang, Xiangxi Zheng, Yilin Wang, Dongxing Mao

arXiv 2608.02109首次发表:更新:

发表机构

Southeast University; Nanjing University; Zhejiang University; National University of Singapore(东南大学; 南京大学; 浙江大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-文本压缩中渲染图像与原生文本表示的跨路径不一致瓶颈,提出自监督对齐框架SPIRAL,通过令牌级OPD与序列级DPO提升模型性能,在VTCBench上显著改善Qwen3-VL-8B的表现并实现域外泛化。

AI 中文摘要

视觉-文本压缩(VTC)将长文本渲染为图像并通过视觉编码器(ViT)编码,将数千个文本令牌压缩为少得多的视觉令牌。然而,由于ViT主要在自然图像上预训练,它捕捉视觉属性(字形、字体大小、布局)而非语言语义,导致渲染图像的表示与原生文本的表示产生偏差。我们将这种跨路径不一致性命名为跨路径不一致,并通过渲染扰动实验表明,它是VTC一个关键却被忽视的瓶颈。我们提出SPIRAL(Self-improving Path Integration and Realignment),这是一种仅利用模型自身文本路径行为作为监督的自监督对齐框架,无需外部教师或额外标注。SPIRAL在两个互补粒度上运行:用于局部保真度的令牌级在线策略蒸馏(OPD),以及用于全局一致性的序列级偏好优化(DPO)。在VTCBench上,SPIRAL将Qwen3-VL-8B的整体得分从35.10提升至54.02,接近原生文本输入的性能(55.60),并优于规模大30倍的模型。两种粒度展现出互补优势:OPD在检索方面表现出色且样本效率高,而DPO在推理和记忆方面更强,且随数据扩展效果更好。SPIRAL的优势还可泛化到域外基准,证实有效的VTC取决于将渲染图像的表示重新对齐到原生文本语义。

英文摘要

Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens into far fewer visual tokens. However, since the ViT is pretrained predominantly on natural images, it captures visual attributes (glyphs, font sizes, layout) rather than linguistic semantics, causing rendered-image representations to diverge from native-text representations. We term this cross-path inconsistency and show, via rendering perturbation experiments, that it is a critical yet overlooked bottleneck of VTC. We propose SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations. SPIRAL operates at two complementary granularities: token-level on-policy distillation (OPD) for local faithfulness, and sequence-level preference optimization (DPO) for global coherence. On VTCBench, SPIRAL improves the overall score of Qwen3-VL-8B from 35.10 to 54.02, approaching the native text-input performance (55.60) and outperforming models up to 30x larger. The two granularities exhibit complementary strengths: OPD excels at retrieval and is sample-efficient, while DPO is stronger on reasoning and memory and scales better with data. SPIRAL's benefits also generalize to out-of-domain benchmarks, confirming that effective VTC hinges on aligning rendered-image representations back to native-text semantics.

CommentsAccepted to ACM Multimedia 2026 (Oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑