发表机构
UMass Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对13位开发者的206个真实会话实验,发现从交互历史提炼的开发者个性化技能改进有限,而汇集自所有开发者的通用技能增益最大且最一致,为编码智能体的个性化策略提供了实证依据。
AI 中文摘要
基于大语言模型(LLM)的智能体已快速从代码补全工具演变为复杂软件工程任务的求解器。随着开发者与编码智能体长期协作,其偏好会通过多次交互显现,可用于调整智能体行为以更好满足个体开发者需求,捕获并复用这些偏好或能减少重复修正、改善开发者-智能体协作。智能体技能提供了一种无需修改模型参数即可传递经验的轻量机制。然而,现有工作主要聚焦于特定任务的技能,从交互历史中提炼的开发者特定技能是否能泛化到未来任务仍不明确。我们提出了一个从交互轨迹中提取可复用开发者偏好的框架,该框架先通过基于规则的引导和基于证据的优化生成个性化技能,再使用可复现的重放框架结合交互式、基于轨迹的LLM人类开发者模拟器对其进行评估。我们对来自13位开发者的206个真实世界开发者-智能体会话开展实验,将个性化技能与无技能、通用技能及其他用户技能基线进行对比。结果显示,个性化技能相较无技能基线仅提供微小且不一致的改进,而汇集自所有开发者的通用技能则实现了最大且最一致的增益。进一步分析表明,当开发者偏好频繁出现时,个性化技能会更有效,尤其是当他们的历史包含多个与未来任务相关的示例时。这些发现为开发者特定个性化何时有效提供了实证见解,并证明广泛可迁移的过程性知识比开发者特定偏好信号更具鲁棒性。
英文摘要
Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks. As developers collaborate with coding agents over time, their preferences emerge through repeated interactions and can be used to adapt agent behavior to better meet individual developers' needs. Capturing and reusing these preferences may reduce repeated corrections and improve developer-agent collaboration. Agent skills provide a lightweight mechanism for transferring experience without modifying model parameters. However, existing work primarily focuses on task-specific skills, and it remains unclear whether developer-specific skills distilled from interaction histories can generalize to future tasks. We propose a framework for extracting reusable developer preferences from interaction traces. It first generates personalized skills through rule-based bootstrapping and evidence-grounded refinement, and then evaluates them using a reproducible replay framework with an interactive, trajectory-conditioned LLM-based human developer simulator. We conduct an experiment on 206 real-world developer-agent sessions from 13 developers and compare personalized skills against no-skill, generic-skill, and other-user-skill baselines. Personalized skills provide small and inconsistent improvements over the no-skill baseline, whereas generic skills pooled across developers achieve the largest and most consistent gains. Further analysis suggests that personalized skills become more effective when developer preferences appear frequently, particularly when their histories contain multiple examples relevant to future tasks. These findings provide empirical insights into when developer-specific personalization is effective and demonstrate that broadly transferable procedural knowledge can be more robust than developer-specific preference signals.
Comments15 pages, 10 figures