arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02582cs.SE

ACEM:面向智能体软件工程的成本估算模型

ACEM: A Cost Estimation Model for Agentic Software Engineering

Mohammad El-Ramly

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体软件工程成本估算问题,本文提出ACEM模型,分解成本为LLM、HITL、基础设施三类,引入RF、CF、HIS构念并映射传统度量,为智能体开发成本估算提供早期方案。

中文摘要 AI 辅助

传统软件成本估算模型(如COCOMO II、功能点、故事点)假设开发工作量主要由设计、编码、测试等环节的人力驱动。而智能体软件工程中,自主AI智能体承担大量实现工作,人类仅聚焦规划、规格说明与验证,这一假设受到挑战。新的成本维度随之产生:智能体动作中的大语言模型(LLM)令牌消耗、人类在环(HITL)监督工作量、智能体编排与工具的基础设施成本。这些成本具有不确定性:相同任务可能消耗不同令牌、遵循不同推理路径、需要不同程度的人工修正,而此类现象在传统开发中并不存在。因此需要新框架衔接标准规模度量与该成本结构。本文提出ACEM(Agentic Cost Estimation Model,智能体成本估算模型),将智能体开发总成本分解为三个可加维度:LLM成本、HITL成本与基础设施成本。ACEM引入三个智能体动态性相关的构念:修订因子(Revision Factor,RF),用于建模因输出拒绝与重试产生的令牌开销;上下文因子(Context Factor,CF),用于捕捉上下文累积带来的令牌消耗增长;以及HITL强度得分(HITL Intensity Score,HIS),一种四级监督分类方案。它还将用例点、故事点与功能点映射至预估令牌消耗,使组织可复用现有项目范围数据进行智能体成本预测。ACEM作为完全规范的模型结构与校准方法被提出,其常数暂为符号形式,需经实证验证。作为早期提案,它邀请研究界通过真实项目数据对模型进行校准、测试与扩展。

英文摘要

Traditional software cost estimation models, such as COCOMO II, Function Points, and Story Points, assume that development effort is primarily driven by human labor in design, coding, and testing. Agentic software engineering, where autonomous AI agents perform substantial implementation work and humans focus on planning, specification, and validation, challenges this assumption. New cost dimensions arise: large language model (LLM) token consumption across agent actions, Human-in-the-Loop (HITL) oversight effort, and infrastructure costs for agent orchestration and tooling. These costs are nondeterministic: identical tasks may consume different tokens, follow divergent reasoning paths, and require varying human correction, phenomena absent in traditional development. A new framework is needed to bridge standard sizing metrics with this cost structure. This paper proposes ACEM (Agentic Cost Estimation Model), which decomposes total agentic development cost into three additive dimensions: LLM, HITL, and infrastructure cost. ACEM introduces three constructs for agentic dynamics: the Revision Factor (RF), modeling token overhead from output rejection and retries; the Context Factor (CF), capturing rising token consumption as context accumulates; and the HITL Intensity Score (HIS), a four-level oversight classification scheme. It further maps Use Case Points, Story Points, and Function Points to estimated token consumption, enabling organizations to reuse existing project-scoping data for agentic cost forecasting. ACEM is presented as a fully specified model structure and calibration methodology, with constants left symbolic pending empirical grounding. As an early-stage proposal, it invites the research community to calibrate, test, and extend the model through real project data.

补充信息

↑