arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教师的少量步骤作用显著:智能体训练后基于策略的数据增强的成本效益

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

Junze Ye, Jiayi Cheng, Miao Lu, Michal Mankowski, Jose Blanchet, Mohsen Bayati

arXiv 2607.04574首次发表:更新:

发表机构

Stanford University; New York University(斯坦福大学; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型智能体,将基于策略的数据构建视为预算分配问题,通过特定方式形式化设计空间,在多个任务中实验表明少量教师步骤在学习者诱导情境下更具成本效益。

AI 中文摘要

对于大语言模型智能体,监督微调不仅关乎教师标签质量,还涉及标签所依赖的交互上下文。纯行为克隆存在训练与测试时上下文不匹配问题。近期工作通过在学生到达的上下文中查询教师解决此问题。我们将基于策略的数据构建视为预算分配问题,通过展开策略、切换时间分布等形式化设计空间,实验表明少量教师步骤更具成本效益。

英文摘要

For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses full teacher demonstrations, creating a mismatch between teacher-induced contexts seen in training and student-induced contexts encountered at test time. Recent work addresses this mismatch by querying a teacher at contexts reached by the student, often with increasingly elaborate filtering of the teacher's continuations. We instead frame on-policy data construction as a budget-allocation problem: under matched supervision resources, should teacher output be spent on more start-to-finish demos, longer continuations, outcome filtering, or broader coverage of learner-induced contexts? We formalize this design space through the rollout policy, switch-time distribution, continuation horizon, filtering rules, and two complementary costs: teacher inference generated before filtering and teacher supervision retained for SFT. Across HotpotQA, ALFWorld, and Terminal-Bench-Dev, bounded unfiltered teacher continuations at learner-induced contexts improve over pure behavioral cloning at matched budgets. On HotpotQA and ALFWorld, where we run the full comparison, few-step continuations match or exceed success-filtered and critical-context-filtered alternatives. Our findings suggest that a few teacher steps, placed at learner-induced contexts, can be a more cost-efficient supervision allocation than longer or more heavily curated teacher completions.

Comments8/28 update: added funding acknowledgements

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑