arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36651cs.CVcs.AI

FocusVTC:自适应分辨率的高效高性能视觉文本压缩

FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

FangZhi Zhong, Xuerui Qiu, Yuqi Pan, Ya Liu, Shaowei Gu, Bo Xu, Guoqi Li

首次发表
浏览论文内容

中文总结 AI 辅助

FocusVTC通过自适应分辨率打破视觉文本压缩的权衡,结合低DPI全局视图与选择性区域增强,在保持多模态能力的同时实现高效压缩与高性能推理。

中文摘要 AI 辅助

大型语言模型中的长上下文推理会带来巨大的计算和内存成本。视觉文本压缩(VTC)通过将文本渲染为图像来减少输入长度,但固定分辨率渲染造成了压缩与性能之间的权衡:低DPI节省了令牌但牺牲了可读性,而高DPI则在无关内容上花费令牌。我们提出了FocusVTC,通过自适应分辨率打破了这一权衡,同时保留了通用的多模态能力。它结合了压缩的低DPI全局视图与选择性区域增强,并将增强视图集成到持续推理中。我们构建了29.4K个高质量的推理-证据定位(REL)思维链示例(REL-CoT),将推理轨迹与页面索引和边界框关联起来。多分辨率REL监督微调(REL-SFT)教会模型定位相关区域,而组相对策略优化则学习何时增强分辨率以及如何使用由此产生的观察结果,无需单独的持续预训练阶段。在RULER v1上,以72 DPI运行时,FocusVTC在2.9倍输入压缩(包括工具观察)下得分87.4,而Glyph在3.0倍输入压缩下得分为57.5。它在LongBench上超越了其文本输入骨干(56.40对55.86),将MRCR宏平均值提高了13.91分,并在VTCBench上达到了51.19的宏平均值。MRCR延迟评估还显示,与Text相比,在线端到端加速了2.79倍。通用多模态能力得以保留,MMMU从65.12增加到66.73,MME从2424.02增加到2457.62。

英文摘要

Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • Shanghai Jiao Tong University(上海交通大学)
  • Zhongguancun Academy(中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑