用于Tokenmaxxing的最佳编程语言:跨编程语言的编码代理行为研究
The Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages
- Northeastern University(东北大学)
- Wellesley College(韦尔斯利学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究跨Python、Java、Rust和OCaml的编码代理行为,通过评估五个模型发现不同语言的令牌消耗有显著差异。分析代理轨迹结构与内容,揭示其在不熟悉语言中的行为特点,指出按语言的令牌效率对多语言代理基准测试和开发有指导意义。
AI中文摘要:
尽管编码代理目前在多种编程语言中都非常有效,但本文首先表明,成本(以令牌计)会因编程语言而有很大差异。我们在Python、Java、Rust和OCaml的编程问题上评估了五个最新模型。我们仔细控制问题难度,并表明令牌消耗存在显著差异且在各模型中一致。为理解原因,我们分析了代理轨迹的结构和内容。首先,重新执行每个中间解决方案并将每个轨迹抽象为测试结果向量序列,然后标记连续解决方案之间的工作。这揭示了代理在不熟悉的语言中反复产生无法编译的解决方案并修改已通过的解决方案。其次,分析轨迹文本,发现代理在代码注释中规划解决方案,不信任提供的测试而青睐自己发明的输入,并通过在Python中进行原型设计来避开不熟悉的目标语言。我们的结果表明,按语言的令牌效率是在对多语言代理进行基准测试和开发时应考虑的指标,对于令牌最大化者来说,是在最昂贵的语言中工作的指南。
英文摘要:
Although coding agents are now very effective in a variety of programming languages, this paper first shows that the cost (in tokens) can very significantly by programming language. We evaluate five recent models on programming problems in Python, Java, Rust, and OCaml. We carefully control for problem difficulty, and show that there can be stark variation in token consumption that is consistent across models. To understand why, we analyze both the structure and content of agent trajectories. First, we re-execute every intermediate solution and abstract each trajectory as a sequence of test-outcome vectors, then label the work between successive solutions. This reveals agents repeatedly producing noncompiling solutions in unfamiliar languages and revising solutions that already pass. Second, we analyze trajectory text, finding that agents plan solutions in code comments, distrust the provided tests in favor of inputs they invent, and sidestep unfamiliar target languages by prototyping in Python. Our results show that by-language token efficiency is a metric that should be considered when benchmarking and developing multilingual agents, and, for the tokenmaxxer, a guide to the most expensive language to work in.