arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越指令遵循:基于技能契约的学习落地技能遵循

Beyond Instruction Following: Learning Grounded Skill-Following with Skill Contracts

Jianghan Shen, Zhenjie Liu, Yue Li, Jie Huang, Siqi Luo, Yiming Cheng, Yizhi Yao, Kaijie Zhang, Cheng Tang, Minghui Zhang, Ming Hu, Yirong Chen, Ziyan Huang

arXiv 2610.05161首次发表:更新:

发表机构

Nanjing University; Shanghai Artificial Intelligence Laboratory; Peking University; Yiyue Technology; Tsinghua University; Northeastern University(南京大学; 上海人工智能实验室; 北京大学; 一岳科技; 清华大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对指令遵循仅关注最终答案的局限,提出基于技能契约的落地技能遵循框架,通过契约运行时提供密集训练信号和反馈,显著提升协议完成率并保持任务结果。

AI 中文摘要

指令遵循通常强制执行离散的、响应级别的需求,而专家编写的技能则规定了跨越多个阶段和环境交互的程序性需求。给定这样的技能,我们训练执行器执行所有必需阶段,而不仅仅关注最终答案。因此,我们引入了落地技能遵循(Grounded Skill-Following),它要求智能体通过将决策落地于环境观察,在其所需阶段中执行一个固定的、专家编写的技能。为了实现可验证的程序性执行,我们将每个技能制定为技能契约,该契约将可见的技能指令与显式的契约运行时相结合。该运行时规定了所需阶段、可允许动作、允许的转换以及接受的终止。这种结构在整个执行过程中提供了密集的、可验证的训练信号。我们利用这一点引入了已验证进度信用(Verified Progress Credit),它在契约里程碑首次完成时分配奖励,并将它们聚合到轨迹回报中以指导策略优化。在推出过程中,契约运行时持续跟踪状态转换以提供契约状态反馈(Contract-State Feedback),该反馈指示最新动作是否被接受,并引导智能体朝向有效的下一动作。为了衡量程序性合规性,我们引入了协议完成率(PCR),定义为通过所有必需阶段达到接受的终止,并将其与最终任务结果(Task Outcome)解耦。与我们的框架联合训练,Qwen3.5-4B在数学任务上达到了99.27%的协议完成率,在搜索任务上达到了99.96%,同时在任务结果上略微优于原始基线(分别为82.95%和46.61%)。对照研究考察了技能指令、训练信号和契约状态反馈如何影响这两个指标,而消融干预则评估了对观察内容的行为依赖性。

英文摘要

Instruction following typically enforces discrete, response-level requirements, whereas an expert-authored skill prescribes procedural requirements spanning multiple phases and environment interactions. Given such a skill, we train the executor to execute all required phases instead of focusing solely on the final answer. We therefore introduce Grounded Skill-Following, which requires an agent to execute a fixed, expert-authored skill across its required phases by grounding decisions in environment observations. To achieve verifiable procedural execution, we formulate each skill as a skill contract combining visible skill instructions with an explicit contract runtime. The runtime specifies required phases, admissible actions, permitted transitions, and accepted termination. This structure provides a dense, verifiable training signal throughout execution. We leverage this by introducing Verified Progress Credit, which assigns rewards upon the initial completion of contract milestones and aggregates them into the trajectory return to guide policy optimization. During rollout, the contract runtime continuously tracks state transitions to provide Contract-State Feedback, which indicates whether the latest action is accepted and guides the agent toward valid next actions. To measure procedural compliance, we introduce the Protocol Completion Rate (PCR), defined as reaching accepted termination through all required phases, and decouple it from the final Task Outcome. Jointly trained with our framework, Qwen3.5-4B achieves Protocol Completion Rates of 99.27% on Math and 99.96% on Search, while slightly outperforming original baselines in Task Outcome (82.95% and 46.61%, respectively). Controlled studies examine how skill instructions, training signals, and contract-state feedback affect both metrics, while withholding interventions evaluate behavioral dependence on observation content.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑