AgentBrew:从原始真实世界轨迹中离线学习工具使用智能体
AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
浏览论文内容
中文总结 AI 辅助
AgentBrew通过回顾性任务推断和PMI信用分配,从原始轨迹离线学习工具使用策略,在三个MCP应用上显著提升性能。
中文摘要 AI 辅助
基于LLM的智能体通过工具使用API越来越多地部署在真实世界应用中,然而针对特定环境训练它们仍然从根本上困难:真实世界应用不提供预定义的任务或验证器,没有忠实的模拟器,并且大规模环境交互的预算有限。在本文中,我们提出AgentBrew,一种离线训练框架,它从单批原始交互轨迹中学习有效的工具使用策略,无需任务验证器或迭代的在线策略回滚。智能体首先探索目标环境以收集原始轨迹语料库,不进行质量过滤。为了从这一嘈杂的语料库中提取训练信号,回顾性任务推断根据每条轨迹的实际结果重构对齐的指令,而基于PMI的信用分配通过点互信息(PMI)将轨迹关于推断指令的总信息分解为加性的每动作信用。这些信用加权策略训练目标,放大信息丰富的动作同时抑制无效动作。在三个真实世界MCP应用(GitHub、Notion、PostgreSQL)上,AgentBrew使Qwen3-32B平均提升+8.7准确率/+9.7分数,超越Qwen3-235B(+2.3/+4.4)并优于拒绝采样(+5.9/+10.3)。这些结果表明,细粒度的离线学习可以从原始轨迹中恢复基于过滤的方法会丢弃的有用监督。代码可在该https URL获取。
英文摘要
LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limited budget for large-scale environment interaction. In this paper, we propose \textbf{AgentBrew}, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts. The agent first explores the target environment to collect a raw trajectory corpus without quality filtering. To extract training signal from this noisy corpus, \emph{retrospective task inference} reconstructs an aligned instruction for each trajectory based on its actual outcome, and \emph{PMI-Based credit assignment} decomposes the trajectory's total information about the inferred instruction into additive per-action credits via pointwise mutual information (PMI). These credits weight the policy training objective, amplifying informative actions while suppressing ineffective ones. On three real-world MCP applications (GitHub, Notion, PostgreSQL), AgentBrew improves Qwen3-32B by +8.7 Acc / +9.7 Score on average, surpassing Qwen3-235B (+2.3 / +4.4) and outperforming rejection sampling (+5.9 / +10.3). These results demonstrate that fine-grained offline learning can recover useful supervision from raw trajectories that filtering-based approaches would discard. The code is available at https://github.com/alphatogo/AgentBrew
发表机构
- Nanyang Technological University(南洋理工大学)
- Kuaishou Technology(快手科技)
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。