发表机构
Infobip(Infobip)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Infobip团队开发的四阶段编码智能体操作工作流,通过前置人力审查、分阶段上下文管理应对故障,同时指出工作流有效性指标缺失等两个未解决问题。
AI 中文摘要
基于大语言模型(LLM)的编码智能体将基础模型与塑造智能体行为的管控框架相结合。对于非平凡任务,从业者如何构建与编码智能体的协作方式,决定了是否能获得可靠结果。我们报告了由Infobip人工智能研究团队开发的操作编码智能体的分阶段工作流,该工作流将智能体辅助开发划分为四个阶段,其中人力投入前置,且随着工件成熟,授权程度逐步提升。上下文管理是核心问题,通过在每个阶段应用四种策略来应对已知故障模式。从从业者经验来看,我们观察到上游研究与规划中的错误会在后续阶段累积,而修正生成的代码可能会引入冗余与脆弱性,这推动了人力审查的前置。我们确定了两个未解决的问题:缺乏衡量工作流有效性的指标,以及形式化的上下文管理组件与从业者所需的工作流级模式之间存在差距。
英文摘要
LLM-based coding agents combine a foundation model with a harness that shapes agent behavior. For non-trivial tasks, how practitioners structure their work with the coding agents determines whether reliable results follow. We report on a phased workflow for operating coding agents developed by the AI research team at Infobip. The workflow structures agent-assisted development into four phases where human effort is front-loaded and delegation increases as artifacts mature. Context management is the central concern, addressed through four strategies applied at each phase to counter known failure modes. From practitioner experience, we observe that upstream errors in research and planning can compound across later phases, while correcting generated code can introduce bloat and fragility. This motivates front-loading human review. We identify two open problems: the absence of metrics for workflow effectiveness and the gap between formalized context management components and the workflow-level patterns that practitioners need.
Comments2 pages, industry paper, to appear in proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), 2026