arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15962cs.CLcs.CV

SEER:基于选择性视觉-文本压缩的长上下文推理

SEER: Long-Context Reasoning via Selective Visual-Text Compression

  • The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • University of Cambridge(剑桥大学)
  • University of Technology Sydney(悉尼科技大学)
  • The University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Jiawei Xu, Zhilin Zhai, Jinrui Fang, Ruohan Xu, Mingfei Lu, Yi Zhang, Guanchu Wang, Tianlong Chen, Ying Ding

中文总结 AI 辅助

SEER是结合视觉压缩效率与文本推理精度的框架,经监督微调后在LongBench等长上下文基准测试中,准确率优于Glyph-9B、Qwen3-8B等基线模型,可提升提取精度并保留提示token节省量。

中文摘要 AI 辅助

对于大语言模型而言,由于文本token注意力的二次复杂度,长上下文推理的计算成本仍然很高。视觉-文本压缩通过将文本渲染为图像并使用视觉-语言模型处理,通常可减少token使用量,是一种很有前景的替代方案。然而,现有方法会应用均匀压缩,不考虑查询相关性,在需要详细提取的地方可能会牺牲精度。我们提出SEER,这是一个通过视觉扫描学习选择与查询相关的图像、仅在需要时检索文本内容的框架,结合了视觉压缩的效率和基于文本推理的精度。通过对工具交互轨迹进行监督微调,SEER学习了用于选择和检索的自适应工具调用。在长上下文基准测试中,SEER通过选择性文本检索提高了提取精度,同时保留了相对于全文基线的平均提示token节省量。在LongBench上,SEER实现了51.11%的平均准确率,超过视觉-文本基线Glyph-9B 2.33个百分点,超过Qwen3-8B 3.49个百分点。代码可在此URL访问

英文摘要

Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images and processing them with vision-language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potentially sacrificing precision where detailed extraction is required. We present SEER, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning. Through supervised fine-tuning on tool-interaction trajectories, SEER learns adaptive tool invocation for selection and retrieval. Experiments on long-context benchmarks show that SEER improves extraction precision through selective text retrieval while retaining average prompt-token savings relative to full-text baselines. On LongBench, SEER achieves 51.11% average accuracy, outperforming the visual-text baseline Glyph-9B by 2.33 points and Qwen3-8B by 3.49 points. Code can be accessed at https://github.com/jiaweixu98/SEER

补充信息

↑