发表机构
University of Oxford; Zhejiang University(牛津大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出规范程序性动作(CPAs)标注协议,通过锚点与证据链接记录工具使用轨迹中的动作,经零售案例验证高可重复性,提供可审计的标注工具。
AI 中文摘要
工具使用智能体轨迹能识别消息和API调用,但程序性分析还需要明确的动作单元及其证据的可检查链接。我们提出规范程序性动作(CPAs),一种标注协议,记录程序性功能、其首个智能体事件锚点、实现该功能的智能体事件以及独立的上下文证据。多个动作可共享一个消息锚点,而无需推断消息内顺序。一项零售案例研究通过开放归纳、记录整合和连续应用审计,生成了一个带版本号的24条目编码手册。两个隔离的LLM上下文在轨迹层面标注了32条与开发不相交的轨迹,分别产生499和491个出现次数,锚点标签重叠度A=0.982。要求相同的上下文事件引用将重叠度降至0.798。这些是结构可重复性度量,而非语义准确性:26个任务ID中有16个也出现在开发中,且历史工具负载被截断至110个字符。回顾性对照显示,合并所有标签将重叠度提升至0.986,而简单端点规则以0.997的重叠度重现了工具锚定部分。助手消息动作的重叠度为0.971,每个标签的最小值为0.816。将冻结的编码手册应用于另外244条轨迹,产生4,058条记录,包括八个诊断结果。贡献在于一个明确、可审计的标注工具及其构建和测量限制的案例研究;人类参考有效性和下游实用性仍有待确立。
英文摘要
Tool-use agent traces identify messages and API calls, but procedural analyses also need explicit units of action and inspectable links to their evidence. We present Canonical Procedural Actions (CPAs), an annotation protocol that records a procedural function, its first agent-event anchor, the agent events that realize it, and separate contextual evidence. Multiple actions may share a message anchor without an inferred within-message order. A retail case study produces a versioned 24-entry codebook through open induction, recorded consolidation, and successive application audits. Two isolated LLM contexts annotate 32 trajectories disjoint from development at the trajectory level, producing 499 and 491 occurrences with anchor-label overlap A=0.982. Requiring identical context-event references reduces overlap to 0.798. These are structural repeatability measures, not semantic accuracy: 16 of 26 task IDs also occur in development, and historical tool payloads were truncated to 110 characters. Retrospective controls show that collapsing all labels raises overlap to 0.986, while simple endpoint rules reproduce the tool-anchored portion with 0.997 overlap. Assistant-message actions have 0.971 overlap, with a per-label minimum of 0.816. Applying the frozen codebook to 244 further trajectories yields 4,058 records, including eight diagnostic outcomes. The contribution is an explicit, auditable annotation instrument and a case study of its construction and measurement limits; human-reference validity and downstream utility remain to be established.
Comments23 pages, 5 figures, 9 tables. Includes ancillary files for reproducing the reported analyses