arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

帕累托占优澄清:通过PPO-Lagrangian预算约束对编码大语言模型进行后训练

Pareto-Dominant Clarification: Post-Training Coding LLMs via PPO-Lagrangian Budget Constraints

Abhinav Rajput, Acey Vogelstein

arXiv 2610.04089首次发表:更新:

发表机构

Center for Data Science; New York University(数据科学中心; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将编码智能体的澄清行为建模为约束马尔可夫决策过程,用PPO-Lagrangian后训练Qwen2.5-Coder,在预算约束下同时提升准确率并降低澄清率,发现未约束系统存在帕累托低效。

AI 中文摘要

在模糊指令或用户提示下运行的编码智能体必须决定是提出澄清问题还是直接尝试解决方案。虽然用户的澄清可能提高智能体解决方案的正确性,但每次来回交互都会产生用户和系统成本,形成明确的准确性与效率权衡。现有工作研究澄清行为,但未在可执行的澄清预算下训练策略;基于惩罚的方法通常需要在所有澄清预算水平上扫描单独的系数调整。我们将澄清问题建模为约束马尔可夫决策过程(CMDP),并使用PPO-Lagrangian对Qwen2.5-Coder-7B-Instruct进行后训练,以优化编码准确性,同时满足预期的提问预算约束。在带有GPT-4o-mini预言机模拟器的HumanEvalComm上评估,所得策略表明未调整的澄清行为是帕累托低效的:受预算约束的策略可以同时实现比基线模型更高的准确性和更低的澄清率。在各个预算水平上,我们观察到对数形状的帕累托前沿,额外澄清的收益递减。收益并非来自简单地整体提出更多问题,而是来自改进的问题定位和模糊条件下的更好代码生成。在没有显式监督的情况下,训练后的策略学会非均匀地分配澄清预算,在更困难(多重退化)任务上更频繁地提问。这些结果表明,无约束的交互式大语言模型系统可能系统地低效使用澄清。

英文摘要

Coding agents operating under ambiguous instructions or user prompts must decide whether to ask clarifying questions or attempt a solution directly. While clarification from the user may improve the correctness of the agent's solution, each back-and-forth interaction incurs user and system costs, forming an explicit accuracy vs. efficiency tradeoff. Existing works study clarification behavior but do not train policies under enforceable clarification budgets; penalty-based approaches typically require separate coefficient tuning swept across all clarification budget levels. We formulate clarification as a Constrained Markov Decision Process (CMDP) and post-train Qwen2.5-Coder-7B-Instruct with PPO-Lagrangian to optimize coding accuracy, subject to an expected question-budget constraint. Evaluated on HumanEvalComm with a GPT-4o-mini oracle simulator, the resulting policies reveal that untuned clarification behavior is Pareto-inefficient: budget-constrained policies can simultaneously achieve higher accuracy and lower clarification rates than the baseline model. Across budget levels, we observe a log-shaped Pareto frontier with diminishing returns to additional clarification. Gains arise not from simply asking more questions overall, but from improved question targeting and better code generation under ambiguity. Without explicit supervision, trained policies learn to allocate clarification budget non-uniformly, asking more frequently on tougher (multi-degradation) tasks. These results suggest that unconstrained interactive LLM systems may systematically use clarification inefficiently.

CommentsAccepted to the LMP Workshop at EMNLP 2026. 17 pages, 8 figures, 1 table. Abhinav Rajput and Acey Vogelstein contributed equally. Code: https://github.com/Abhinav0710rajput/coding_llms_cmdp

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑