arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码智能体能通过渲染代码解决仓库级问题吗?一项关于视觉表示的探索性研究

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

Weijie Liang, Yuanfeng Song, Xing Chen, Caleb Chen Cao, Sirui Han, Yike Guo

arXiv 2608.09268首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; ByteDance(香港科技大学; 字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探索将渲染代码作为编码智能体的操作上下文,基于 SWE-bench Verified 实验发现其可降低提示 token 成本但受模型架构限制,仅在原始代码读取为瓶颈时有用,是可行但有条件的压缩机制。

AI 中文摘要

视觉模态近期被探索作为压缩文本 token 的一种方式,包括将代码渲染为图像用于静态代码理解。我们研究这种表示是否可作为智能体式编码的操作上下文,其中智能体必须导航仓库、编辑源文件并验证可执行补丁。使用 SWE-bench Verified,我们在仓库级修复工作流中评估渲染代码,并引入受控智能体设置以区分无引导的仓库探索与更结构化的修复阶段。结果显示情况复杂:渲染代码持续降低提示 token 成本,但节省量不随名义视觉压缩率线性增长;它在很大程度上保留端到端修复准确率,但无法克服底层模型或智能体架构的性能限制,且在激进压缩下可能变得不稳定。进一步分析表明,当原始代码读取是主要瓶颈时,视觉代码最有用;一旦仓库定位结构化,剩余大部分成本来自补丁-测试的试错,视觉压缩的作用有限。总体而言,本研究将渲染代码定位为现实编码智能体的可行但有条件的压缩机制。

英文摘要

Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representation can serve as operational context for agentic coding, where an agent must navigate repositories, edit source files, and verify executable patches. Using SWE-bench Verified, we evaluate rendered code in repository-level repair workflows and introduce controlled agent settings to separate unguided repository exploration from more structured repair stages. Our results show a mixed picture. Rendered code consistently reduces prompt-token cost, but the savings do not increase linearly with the nominal visual compression ratio. It largely preserves end-to-end repair accuracy, but does not overcome the performance limits of the underlying model or agent architecture, and can become unstable under aggressive compression. Further analysis suggests that visual code is most useful when raw source reading is a major bottleneck; once repository localization is structured, much of the remaining cost comes from patch--test trial-and-error, where visual compression has limited leverage. Overall, our study positions rendered code as a viable but conditional compression mechanism for realistic coding agents.

Comments8 pages of main content

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑