发表机构
Alibaba Cloud(阿里云)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出可执行代码知识(ECK)及其单元ECKU,构建Python原型验证其在编码智能体任务中的效果,支持源绑定、验证等功能,为AI编码智能体提供了混合架构方案。
AI 中文摘要
AI编码智能体所需的不止是相关代码片段,还需要业务语义、验证证据、关系以及上下文为当前状态的保证。现有系统通常通过检索、摘要、图谱、规则或反向规范来推断或外部化这些知识。我们研究一种互补表示,其中选定的代码单元直接承载智能体可用的知识。我们提出可执行代码知识(Executable Code Knowledge,ECK),并将可执行代码知识单元(Executable Code Knowledge Unit,ECKU)定义为结合稳定身份、语义、可执行行为、契约、证据、关系、来源、验证状态和查询接口的源绑定对象。我们的Python原型支持代码本地创作、清单导出、证据执行、精确变更行影响、新鲜度检查以及面向智能体的投影。在3个真实Python仓库和26个受控补丁任务中,直接ECK为11/11项带证据的任务提供了可执行测试覆盖,为9/11项提供了精确选择器;隐藏声明的证据将精确恢复率降至1/11(配对精确McNemar检验p=0.0078)。ECK衍生的规则恢复了11/11项精确选择器,表明规则是有效的交付制品,而ECK提供源绑定、验证状态、影响和新鲜度。精确变更行影响与全部26个补丁(12个单元链接)的独立编写标签匹配,精确率、召回率和F1值均为1.000。基于AST的指纹正确分类了50个正向变更和17个无关同文件对照,而静态规则快照未检测到50个陈旧案例中的任何一个。基于模型的补丁审查和跨层研究衡量投影保真度而非独立影响发现。这些结果支持一种混合架构:检索用于覆盖,ECK用于源和证据治理,投影用于交付。
英文摘要
AI coding agents need more than relevant snippets: they need business semantics, validation evidence, relations, and assurance that their context is current. Existing systems usually infer or externalize this knowledge through retrieval, summaries, graphs, rules, or reverse specifications. We investigate a complementary representation in which selected code units directly carry agent-usable knowledge. We introduce Executable Code Knowledge (ECK) and define an Executable Code Knowledge Unit (ECKU) as a source-bound object combining stable identity, semantics, executable behavior, contracts, evidence, relations, provenance, validation state, and a query interface. Our Python prototype supports code-local authoring, manifest export, evidence execution, exact changed-line impact, freshness checking, and agent-facing projections. Across three real Python repositories and 26 controlled patch tasks, direct ECK provides executable test coverage for 11/11 evidence-bearing tasks and exact selectors for 9/11; hiding declared evidence reduces exact recovery to 1/11 (paired exact McNemar p=0.0078). ECK-derived rules recover 11/11 exact selectors, showing that rules are effective delivery artifacts while ECK supplies source binding, validation state, impact, and freshness. Exact changed-line impact matches independently authored labels on all 26 patches (12 unit links; precision, recall, and F1 all 1.000). AST-bounded fingerprints classify 50 positive changes and 17 unrelated same-file controls correctly, whereas static rules snapshots detect none of the 50 stale cases. Model-backed patch-review and cross-layer studies measure projection fidelity rather than independent impact discovery. These results support a hybrid architecture: retrieval for coverage, ECK for source and evidence governance, and projections for delivery.
Comments11 pages. Submitted to AgenticDev 2026, co-located with ASE 2026