arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自分割轨迹:智能体声明边界作为训练单元

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

Jingxi Wei

arXiv 2608.02302首次发表:更新:

AI 中文总结

该研究提出智能体声明边界的采集时语义自分割方法,将其作为训练单元,实验表明该分段具有一致性,下游DPO在特定任务中表现优于对照组。

AI 中文摘要

长程编码智能体轨迹与用于训练的信用单元匹配度较差:单个动作无稳定价值,回合标签将有效探索与放弃方向合并,固定窗口在日志记录机制失效处截断。我们提出采集时语义自分割,即声明式契约让执行智能体在生成轨迹时暴露自身边界,该机制通过可证伪因果假设实例化,连续采纳会呈现变长语义阶段,无需里程碑词汇、黄金补丁、环境回放、教师逻辑或回溯分段器来设置边界。由于智能体命名其猜想,评审者可按名称否定,使该协议能产生记录工作罕见的“先因错误后修正”转换;一次采集即可生成四个监督目标,包括回合标签丢弃的恰好失败区域的审计监督。随后我们探究删除声明后仍留存的内容:仅给定分界点而非假设时,模型将动作块归因于其主导假设的准确率超随机水平两倍,优于相同轨迹上的等长块(配对符号检验 p=0.0002),通过词汇控制且在标签置换下失效。要求放置边界时,代码盲标注器在40个案例中匹配24个,随机放置匹配11.5个,机械测试事件规则在严格到宽松的扫描两端均未优于随机,因此这些分段具有一致性且难以廉价复制。下游方面,对2551个阶段边界对的DPO(直接偏好优化)在91个对抗保留项上未改变任何决策,60个匹配结构项中有4个改变,均为错改对,两个对照组无改变:来自一个生成器的1825对显示,需调整的变量是语料库多样性而非边界。

英文摘要

Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall. We introduce collection-time semantic self-segmentation, in which a declarative contract has the acting agent expose its own boundaries while the trajectory is generated. Instantiated with falsifiable causal hypotheses, successive adoptions expose variable-length semantic phases, and no milestone vocabulary, gold patch, environment replay, teacher logits, or retrospective segmenter places a boundary. Because the agent names its conjecture, a reviewer can negate it by name, which lets our protocol manufacture wrong-cause-then-correction transitions that recorded work rarely contains; one collection then yields four supervised targets, including audit supervision from exactly the failed regions an episode label discards. We then ask what survives deleting the declaration. Given the cut points but not the hypothesis, a model attributes action blocks to their governing hypothesis at over twice chance, beating equal-length blocks over the same trajectories (paired sign test $p = 0.0002$), surviving a lexical control and collapsing under label permutation. Asked instead to place boundaries, a code-blind annotator matches 24 of 40 where random placement matches 11.5, while a mechanical test-event rule beats chance at neither end of a strict-to-permissive sweep. The segments are therefore coherent and not cheaply reproducible. Downstream, DPO on 2,551 phase-boundary pairs changes no decision on 91 adversarial held-out items, while four of 60 change on matched-construction items, all wrong to right, where two controls change none: with 1,825 pairs from one generator, the variable to vary next is corpus diversity, not the boundary.

Comments20 pages, 6 figures, 11 tables. Includes appendices with full controls and ablations

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑