arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16096cs.SEcs.AI

教练 Qwen3 Coder 30B 像 CodeClash 竞技场智能体一样思考

Coaching Qwen3 Coder 30B to Think Like a CodeClash Arena Agent

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Ivy Ning Zhang

AI总结:

针对弱代码智能体在长时程竞技场交互中的失败,提出 ReAct SFT 和轨迹质量加权 SFT 改进 Qwen3-Coder-30B,使其在 CodeClash 锦标赛中超越 Qwen3 Coder Plus。

AI中文摘要:

大型语言模型编码智能体最近在软件任务中变得有用,但较弱的或开放权重的智能体在可靠地解释用户意图和执行复杂的多步骤工作流方面仍然困难。这种差距在长时程设置中尤为明显,在这种设置中,智能体必须在交互约束下反复检查先前的结果、诊断失败并选择下一个代码编辑。这引出了一个自然的问题:我们能做些什么来改进弱代码智能体的思考过程?我们在 CodeClash 中研究这个问题,这是一个代码竞技场基准,原始工作通过多轮锦标赛评估了 6 个竞技场中的 8 个商业编码智能体。由于 Qwen3 Coder Plus 在其中排名最后,我们以开放权重的 Qwen3-Coder-30B 作为案例研究,并调查如何利用来自更强智能体的蒸馏知识来改进它。我们的分析表明,Qwen3-Coder-30B 并未针对竞技场式交互进行良好优化:它经常产生语法和协议破坏性错误,并表现出跨轮次的弱策略适应性。这些失败仅靠普通指令微调难以纠正,因为离线 SFT 无法直接验证生成的动作是否有效或有益。为了解决这个问题,我们提出了 ReAct SFT,它将教师轨迹重写为显式的 [obs][thought][act] 链,以及轨迹质量加权 SFT,它重新加权样本以鼓励编辑后检查。ReAct SFT 显著改善了策略行为,我们的微调模型在锦标赛评估中优于原始的 Qwen3 Coder Plus。

英文摘要:

Large language model coding agents have recently become useful for software tasks, but weaker or open-weight agents still struggle to reliably interpret user intent and execute complex multi-step workflows. This gap is especially visible in long-horizon settings, where an agent must repeatedly inspect prior outcomes, diagnose failure, and choose the next code edit under interaction constraints. It motivates a natural question: what can we do to improve the thinking process of a weak code agent? We study this question in CodeClash, a code-arena benchmark where the original work evaluates 8 commercial coding agents across 6 arenas through multi-round tournaments. Since Qwen3 Coder Plus ranks last among them, we take the open-weight Qwen3-Coder-30B as a case study and investigate how to improve it with distilled knowledge from stronger agents. Our analysis shows that Qwen3-Coder-30B is not well optimized for arena-style interaction: it frequently produces syntax and protocol-breaking errors and exhibits weak strategic adaptation across rounds. These failures are difficult to correct with vanilla instruction tuning alone, since offline SFT cannot directly verify whether a generated action is valid or beneficial. To address this, we propose ReAct SFT, which rewrites teacher trajectories into explicit [obs][thought][act] chains, and trajectoryquality weighted SFT, which reweights samples to encourage post-edit checking. ReAct SFT substantially improves strategic behavior, and our fine-tuned model outperforms the original Qwen3 Coder Plus in tournament evaluation.

↑