发表机构
Florida Institute of Technology(佛罗里达理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TraceCoder通过三种机制实现可解释可审计代码生成,在30项算法编程任务上Mean Chg%达30%,使自动代码生成的内部叙事可审计复现,保障生产部署的信任与问责。
AI 中文摘要
当代基于大语言模型(LLM)的编码智能体生成的代码如同黑箱输出:每行代码背后的原理被隐藏,通过基准驱动修复实现的代码演化是短暂的,事后审计也无法进行。本文提出一种代码生成概念,通过三种互补机制解决这些缺陷:(i)关系型代码片段历史模式,每条修复事件记录基准参考、轮次编号、失败文本和LLM解释,支持完整溯源查询;(ii)基于浏览器的可视化工具,将此历史渲染为带有热图和悬停注释的源代码;(iii)带有树节点分隔符的竞争性分数位置键索引方案,为每个代码片段分配稳定的字典序标识符,在不干扰周围代码行的情况下实现细粒度跟踪。我们在30项算法编程任务上评估TraceCoder,涵盖字符串处理、数学计算和数据结构操作,涉及两种提供商配置。其中10项任务因存在微妙边界行为,耗尽了6次迭代预算。平均变化百分比(Mean Chg%)达到30%,在20任务子集上,使用Gemini 2.0 Flash作为唯一提供商时,30%的代码片段带有可溯源的修复事件行,而该比例为21%。三项详细案例研究展示了系统如何解释哪些特定基准故障塑造了最终程序的每一行。所提出的机制使自动代码生成的内部“叙事”可审计和可复现,这是生产部署中信任与问责制的必要属性。
英文摘要
Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records, per repair event, the benchmark reference, round number, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualisation tool that renders this history as heat-mapped, hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-ordered identifiers to each code snippet, enabling fine-grained tracking without disrupting surrounding lines. We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations. Of these, 10 exhaust the 6-iteration budget on tasks with subtle edge-case behaviour. Mean Chg% reaches 30%, three in ten code snippets carry a traceable repair-event row, compared to 21% when using Gemini 2.0 Flash as sole provider on a 20-task subset. Three detailed case studies demonstrate how the system explains which specific benchmark failures shaped each line of the final program. The proposed mechanism makes the internal "narrative" of automated code generation auditable and replayable, a property essential for trust and accountability in production deployments.
CommentsSubmitted Version (version submitted on May deadline to AGENTICS 2026). Version of Record to appear in AGENTICS proceedings