arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TraceCompiler:技能引导的LLM智能体轨迹挖掘与编译为近确定性工作流

TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows

Salma El Yadouni, Guanyi Li

arXiv 2608.02680首次发表:更新:

发表机构

EPFL; Binome Technologies(洛桑联邦理工学院; Binome科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TraceCompiler系统,通过技能引导挖掘含噪声的LLM智能体轨迹并编译为近确定性工作流,在T1、AppWorld数据集上验证了依赖恢复的高精度,实现了Venmo资金请求意图的调用减少。

AI 中文摘要

使用工具的语言模型智能体反复重新执行已完成的流程,生成的轨迹混合了可复用结构与重试、探索、偶然排序及重复查找操作。本文提出TraceCompiler,这是一个技能引导的系统,可挖掘含噪声的智能体轨迹集群并将其编译为可执行的近确定性工作流。该系统仅在消费者参数包含唯一归属至更早生产者的值时,才允许工具间依赖;每条硬边均附带可审计的证据元组,模糊关系则标记为可疑且不施加排序约束。绑定被分类为常量、用户输入、复制输出、变换或残留LLM决策。在T1数据集(规则的机械化形式)上,针对其训练拆分的15775条定义-使用边,该方法恢复生产者-消费者依赖的精度达0.928、召回率达0.943,而相同数据上的邻接方法F1值为0.711、频率阈值直接跟随方法F1值为0.712;编译器技能盲运行在其中250条边上达到0.992的表现。在AppWorld数据集上,我们在确定性模拟器中重放已发布轨迹以恢复被掩盖的返回值,针对563条标记边,该规则达到0.993的精度——这是自一致性检查,因为重放通过相关启发式注入标记。我们编译了两种常见意图:Venmo资金请求意图将34次观测API调用减少至11次运行时调用,在针对基准自身状态测试的留一法执行下,21次测试通过15次,失败折因所需分支从未被观测到而升级(而非执行);Spotify/Todoist意图被编译器正确拒绝编译,因为不可逆副作用欠确定。我们仅测量调用减少量,未测量离线编译成本,因此不宣称净效率结果。

英文摘要

Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups. We present TraceCompiler, a skill-guided system that mines clusters of noisy agent traces and compiles them into executable, mostly deterministic workflows. It admits an inter-tool dependency only when a consumer argument contains a value attributable uniquely to an earlier producer; every hard edge carries an auditable evidence tuple, and ambiguous relations are marked suspected and impose no ordering constraint. Bindings are classified as constants, user inputs, copied outputs, transforms, or residual LLM decisions. On T1, a mechanized form of the rule recovers producer-consumer dependencies at 0.928 precision and 0.943 recall over 15,775 def-use edges of its training split, against 0.711 F1 for adjacency and 0.712 for a frequency-thresholded directly-follows measure on identical data; the compiler skill run blind reaches 0.992 on 250 of those edges. On AppWorld we replay released trajectories in the deterministic simulator to recover masked return values and measure the rule against 563 token edges at 0.993 precision - a self-consistency check, since replay injects tokens by a related heuristic. We compile two recurring intents: a Venmo money-request intent reduces 34 observed API calls to 11 runtime calls and, under leave-one-out execution against the benchmark's own state tests, passes 15 of 21, the failing fold escalating rather than acting because its required branch was never observed; and a Spotify/Todoist intent the compiler correctly refuses to compile, because an irreversible side effect is under-determined. We measure call reduction but not offline compilation cost, so we claim no net efficiency result.

Comments17 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑