更少澄清,更优代码:评测编码助手的跨会话个性化歧义适应能力
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
浏览论文内容
中文总结 AI 辅助
该研究提出个性化歧义适应新任务,构建CAPA评测基准,评估12个LLM在有无用户历史下的表现,提出同用户历史门控方法,助力开发减少重复澄清的长期编码助手。
中文摘要 AI 辅助
AI辅助编码日益将非正式用户意图转化为可执行软件,但编码请求常包含用户特定的、跨任务和会话重复出现的歧义。现有歧义消解方法通常在当前编码会话中孤立处理每个歧义请求,常通过引出额外澄清来解决。然而,同一用户已解决的会话历史是否可作为记忆,用于解决新会话中重复出现的个性化歧义,这一问题尚未得到充分探索。我们将个性化歧义适应定义为一项新任务:给定用户先前已解决的编码会话和新的歧义请求,助手应识别重复出现的歧义模式,生成预期的可执行解决方案,并尽量减少澄清。为评测该任务,我们引入CAPA,它通过六种机制表征个性化编码歧义,并使用受控的三阶段生成流程将这些机制注入无歧义的可执行任务中。CAPA包含60个平衡的用户-歧义单元对应的600个编码会话,其中包括300个保留的评估会话。我们在无历史和同用户历史条件下,使用可执行成功率、首轮成功率和完成轮数对12个近期大语言模型(LLM)进行评测。我们的分析考察了任务难度、用户身份和基于记忆的历史使用情况,并进一步提出同用户历史门控作为一种轻量级推理时方法。CAPA为开发长期编码助手提供了基础,这类助手能更好地使生成的代码与用户意图对齐,同时减少重复澄清。
英文摘要
AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a newly opened session remains underexplored. We formulate personalized ambiguity adaptation as a new task: given a user's previously resolved coding sessions and a new ambiguous request, an assistant should identify the recurring ambiguity pattern, produce the intended executable solution, and minimize clarification. To benchmark this task, we introduce CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects these mechanisms into unambiguous executable tasks using a controlled three-stage generation pipeline. CAPA contains 600 coding sessions across 60 balanced user--ambiguity cells, including 300 held-out evaluation sessions. We evaluate 12 recent LLMs under no-history and same-user-history conditions using executable success, first-turn success, and turns-to-completion. Our analyses examine task difficulty, user identity, and memory-based history use, and we further propose same-user history gating as a lightweight inference-time method. CAPA provides a foundation for developing long-term coding assistants that better align generated code with user intent while reducing repeated clarification.