发表机构
Stanford University; New York University(斯坦福大学; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型智能体,将基于策略的数据构建视为预算分配问题,通过特定方式形式化设计空间,在多个任务中实验表明少量教师步骤在学习者诱导情境下更具成本效益。
AI 中文摘要
对于大语言模型智能体,监督微调不仅关乎教师标签质量,还涉及标签所依赖的交互上下文。纯行为克隆存在训练与测试时上下文不匹配问题。近期工作通过在学生到达的上下文中查询教师解决此问题。我们将基于策略的数据构建视为预算分配问题,通过展开策略、切换时间分布等形式化设计空间,实验表明少量教师步骤更具成本效益。
英文摘要
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses full teacher demonstrations, creating a mismatch between teacher-induced contexts seen in training and student-induced contexts encountered at test time. Recent work addresses this mismatch by querying a teacher at contexts reached by the student, often with increasingly elaborate filtering of the teacher's continuations. We instead frame on-policy data construction as a budget-allocation problem: under matched supervision resources, should teacher output be spent on more start-to-finish demos, longer continuations, outcome filtering, or broader coverage of learner-induced contexts? We formalize this design space through the rollout policy, switch-time distribution, continuation horizon, filtering rules, and two complementary costs: teacher inference generated before filtering and teacher supervision retained for SFT. Across HotpotQA, ALFWorld, and Terminal-Bench-Dev, bounded unfiltered teacher continuations at learner-induced contexts improve over pure behavioral cloning at matched budgets. On HotpotQA and ALFWorld, where we run the full comparison, few-step continuations match or exceed success-filtered and critical-context-filtered alternatives. Our findings suggest that a few teacher steps, placed at learner-induced contexts, can be a more cost-efficient supervision allocation than longer or more heavily curated teacher completions.
Comments8/28 update: added funding acknowledgements