arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像素换程序?关于将源代码作为文本和图像的输入令牌计费的跨供应商案例研究

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

Ronak Bhalgami

arXiv 2607.21672首次发表:更新:

AI 中文总结

研究将源代码作为文本和图像时商业API如何计费,通过可重复案例研究,涵盖五种编程语言、多种源长度及多个模型别名,得出不同模型的图像与文本比率及收支平衡行为差异,还发布了重现研究所需的资料。

AI 中文摘要

长源代码上下文会消耗大量文本令牌,这促使人们提议将代码渲染为图像供视觉语言模型使用。最近的工作探讨了模型在这种转换后是否仍能解决代码任务。我们研究了一个不同的系统问题:商业应用程序编程接口如何计算由此产生的请求。我们针对原始源文本和紧凑的渲染图像表示,给出了供应商报告的输入令牌的可重复测量案例研究。该基准在五种编程语言、20到2000行的九种源长度以及Anthropic、OpenAI和谷歌Vertex AI公开的15种可用模型别名之间配对请求。这些别名大致可归结为五种不同的计费签名,并非独立的模型复制。在675个完整的文本/图像对中,总的图像与文本比率分别为0.135、0.194和0.242,分别对应报告的输入令牌减少86.5%、80.6%和75.8%。这些总数掩盖了实际不同的收支平衡行为:Anthropic和OpenAI的图像在每个测试大小下的计数较低,而Gemini图像在20行时需要的令牌数是文本的6.95倍,总体上仅在200行时超过文本。有针对性的审计还重现了Gemini图像在页面边界上的非单调计费情况。本研究测量了一个紧凑渲染管道的黑盒请求计费。它没有测量语义保真度、任务准确性、延迟、货币成本或编码代理效率。我们发布了重现和扩展该研究所需的脚本、修订固定的语料库规范、原始使用记录、验证器和确定性分析。

英文摘要

Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whether models can still solve code tasks after this transformation. We examine a different systems question: how commercial APIs count the resulting requests. We present a reproducible measurement case study of provider-reported input tokens for raw source text and a compact rendered-image representation. The benchmark pairs requests across five programming languages, nine source lengths from 20 to 2,000 lines, and 15 available model aliases exposed by Anthropic, OpenAI, and Google Vertex AI. These aliases collapse to approximately five distinct accounting signatures and are not independent model replications. Across 675 complete text/image pairs, aggregate image-to-text ratios are 0.135, 0.194, and 0.242, corresponding to reported input-token reductions of 86.5\%, 80.6\%, and 75.8\%, respectively. These totals conceal materially different break-even behavior: Anthropic and OpenAI images receive lower counts at every tested size, while Gemini images require 6.95 times as many tokens at 20 lines and cross below text only at 200 lines in the aggregate. A targeted audit also reproduces non-monotonic Gemini image accounting across a page boundary. This study measures black-box request accounting for one compact rendering pipeline. It does not measure semantic fidelity, task accuracy, latency, monetary cost, or coding-agent efficiency. We release the scripts, revision-pinned corpus specification, raw usage records, validators, and deterministic analysis needed to reproduce and extend the study.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑