arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CodeShrink:面向高效多模态代码理解的自适应视觉压缩

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Zipeng Xie, Junliang Liu, Wang Luo, Yuan Jiang, Yintong Huo, Michael Lyu

arXiv 2607.29637首次发表:更新:

AI 中文总结

CodeShrink是含无空白渲染、自适应压缩配置、主导令牌选择的自适应视觉压缩框架,可减少代码理解的视觉令牌用量,在三类代码任务上优于基线。

AI 中文摘要

将源代码渲染为图像为降低多模态大语言模型(MLLMs)的输入成本提供了可行途径。调整图像分辨率可在视觉令牌成本与内容保真度之间进行权衡,但仅缩放分辨率会忽略两类低效来源:换行和缩进产生的空白区域,以及与当前指令无关的代码区域。此外,最佳压缩设置因输入、任务和模型而异,限制了固定比例策略的适用性。我们提出CodeShrink,这是一个包含三个组件的自适应视觉压缩框架:无空白渲染将依赖空白的布局替换为紧凑布局和显式结构标记,消除布局诱导的令牌;自适应压缩配置使用经强化学习训练的轻量智能体,为每个输入预测平衡令牌效率与可读性的设置;主导令牌选择则在推理阶段结合指令与代码图像,修剪任务无关的视觉令牌。我们在代码问答、克隆检测和代码补全任务上评估CodeShrink,结果显示其可减少高达71.2%的视觉令牌使用量,同时性能匹配或优于未压缩的纯文本输入,且在所有三个任务中均持续优于基于文本和视觉压缩的基线。这些结果表明,结合布局压缩、自适应配置与指令感知剪枝可提升多模态代码理解的效率,代码已公开。

英文摘要

Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost against content fidelity. However, resolution scaling alone overlooks two sources of inefficiency: blank regions created by line breaks and indentation, and code regions irrelevant to the current instruction. Moreover, the best compression setting varies across inputs, tasks, and models, limiting fixed-ratio strategies. We propose CodeShrink, an adaptive visual compression framework with three components. Blank-Free Rendering replaces whitespace-dependent layouts with compact layouts and explicit structural markers, removing layout-induced tokens. Adaptive Compression Configuration uses a lightweight agent trained with reinforcement learning to predict a per-input setting that balances token efficiency and readability. Dominant Token Selection jointly analyzes the instruction and code image to prune task-irrelevant visual tokens during inference. We evaluate CodeShrink on code question answering, clone detection, and code completion. CodeShrink reduces visual token use by up to 71.2\% while matching or exceeding uncompressed text-only inputs, and consistently outperforms text-based and visual compression baselines across all three tasks. These results show that combining layout compaction, adaptive configuration, and instruction-aware pruning can make multimodal code understanding more efficient. Our code is available at https://github.com/vinsontang1/CodeShrink.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑