arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于账本控制的零样本自编排提升大语言模型的代码生成性能

GVS5H: Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

Victor Gao, Vida Khosrowshahi, Ali Khosrowshahi, Xihao Sun, Juhyun Lee, Simon, Lee

arXiv 2608.26480首次发表:更新:

发表机构

Persis Capital Inc.(珀西斯资本公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于账本控制的管理者-工作者零样本自编排架构,在9个模型的100个困难代码问题上验证其可提升大语言模型代码生成性能,且成本低于改用更大模型。

AI 中文摘要

多智能体大语言模型系统被广泛报道优于单模型基线,但相关证据并不一致,且对比通常存在混淆:流水线同时改变令牌预算、工具调用和提示,因此总体收益很少能揭示真正起作用的因素。我们研究了在共享文件系统工作区中引入管理者-工作者架构的效果,该架构无需训练,也无需针对每个基准进行调优,与同一模型单次运行的表现进行对比。在9个模型——5个开放权重模型(参数规模从9B到约2.8T)和4个前沿闭源模型——的100个最新的困难LiveCodeBench问题上,该架构的收益是真实的但具有条件性:对部分模型而言,收益显著且具有统计学意义(Qwen3.8-27B收益+23.4,GPT-5.6-Luna收益+10.6,GPT-5.6-Terra收益+8.0,均基于5次配对运行;Kimi-K3在关闭推理时5次配对运行收益+30.4,p<10^-4,单次运行在128k令牌上限下收益+42;Minimax-M3在关闭推理时5次配对运行收益+11.0,p<10^-4,单次运行在128k令牌上限下收益+12),而对其他模型则收益为零或为负(Qwen3.6-35B在关闭推理时收益为-1至-9)。在本研究中,Opus-5在单次运行中达到最高得分91%。运行管理者大致会使令牌成本增加两倍,但它获得的精度提升比改用更大规模模型的成本更低:带管理者的GPT-5.6-Terra的精度接近Fable 5的单次调用精度(85.0对比87.4,p=0.59),但成本仅为五分之一(每100个问题单次运行成本11.71美元对比61.11美元,p<10^-4);Qwen-27B的对应成本为51.75美元,且任何人都可自行托管其权重。我们的转录分析发现了收益背后的几种机制,其中两种反复出现:上下文管理(即简短的工作者调用和共享笔记组织状态并减少截断)和问题分解。对于开启推理的大模型,提升幅度较小;对于部分关闭推理的模型以及开启推理的较小模型,提升幅度更大。

英文摘要

Frontier coding performance is typically attained with large, costly proprietary models. We introduce ledger-based zero-shot self-orchestration (GVS5H), a training-free method in which fresh instances of one model decompose problems and coordinate through a shared file system. Across eleven open and closed-weight models on the 100 latest hard LiveCodeBench problems, the method yields as much as 25.6 points improvement, boosting several cheaper models to frontier-level performance. Orchestrated Qwen3.8 Flash Next scores 93.0% against Fable 5's 90.4% at 9% of the cost, while the smaller Qwen3.8-27B reaches 92.4%. Gains are not universal: some models are unchanged or worse. Transcript analysis attributes the gain to decomposition and persistent context. Inference-time organization can reach or exceed frontier coding accuracy at a fraction of the cost on self-hostable weights.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑