arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18959cs.CL

LangSelect:面向LLM代码生成的成本感知目标语言路由

LangSelect: Cost-Aware Target-Language Routing for LLM Code Generation

Son Ha Xuan, Phat T. Tran-Truong, Xuan-Bach Le, Nghia Duong-Trung

首次发表
浏览论文内容

中文总结 AI 辅助

针对多语言可验证的代码生成任务,提出验证感知路由器LangSelect,在生成前选择目标语言并在失败时回退,实验表明可显著降低令牌成本,同时保持高通过率,定义了成本-正确性权衡前沿。

中文摘要 AI 辅助

LLM代码生成系统通常在解码前选择目标编程语言,并将该选择视为固定不变。我们表明,对于语言灵活的编程任务——即多个目标语言可接受且可通过相同测试验证的任务——这一选择是一个可衡量的成本杠杆:同一任务的已验证实现,其生成的令牌长度可能差异显著。我们引入LangSelect,一个验证感知的路由器,在生成前选择目标语言,并在首次尝试失败时回退。为分离离线路由机会与端到端行为,我们评估了已验证解决方案重放(从已接受的语料库解决方案中选择)和实时GPT-5生成(每次生成尝试均计费,包括失败和回退)。在MultiLang-Bench(一个包含3000个任务、8种语言的已验证语料库)上,重放显示了显著的语言路由空间。在450个保留任务的实时评估中,训练集划分的Domain启发式基线将包含包装器和入口点开销的harness代理令牌减少了50.3%,回退后通过率为92.9%;而学习的CodeBERT+元数据选择器在回退后达到最高通过率93.8%,令牌增加3.7%。这些结果表明,输出语言路由可以为单元测试可验证的代码生成定义一个实用的成本-正确性前沿。

英文摘要

LLM code-generation systems usually choose a target programming language before decoding and treat that choice as fixed. We show that, for language-flexible programming tasks -- tasks where several target languages are acceptable and checkable by the same tests -- this choice is a measurable cost lever: verified implementations of the same task can differ substantially in generated-token length. We introduce LangSelect, a verification-aware router that selects the target language before generation and falls back when the first attempt fails. To separate offline routing opportunity from end-to-end behavior, we evaluate verified-solution replay, which chooses among already accepted corpus solutions, and live GPT-5 generation, which charges every generation attempt, including failures and fallbacks. On MultiLang-Bench, a 3,000-task, 8-language verified corpus, replay shows substantial language-routing headroom. In live evaluation on 450 held-out tasks, a train-split Domain heuristic baseline reduces harness-proxy tokens, which include wrapper and entrypoint overhead, by 50.3\% at 92.9\% pass after fallback, while a learned CodeBERT+metadata selector reaches the highest pass after fallback, 93.8\%, with a 3.7\% token increase. These results show that output-language routing can define a practical cost-correctness frontier for unit-test-verifiable code generation.

发表机构

  • RMIT University(皇家墨尔本理工大学)
  • Ho Chi Minh City University of Technology (HCMUT), VNU-HCM(越南国家大学胡志明市分校理工大学)
  • German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑