发表机构
Sahara AI; University of Southern California(撒哈拉人工智能公司; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有长时程智能体研究忽略环境动态性的问题,提出ReLiveGym环境,在数周重放的真实数据流中评估智能体,发现行动时机机制是重要设计轴,并验证了反馈学习的作用。
AI 中文摘要
随着大语言模型(LLM)智能体被广泛采用,它们越来越多地被部署到需要持续监控或周期性动作(如市场分析)的任务中。这些智能体被期望在无人值守的情况下运行数天或数周,在正确的时机采取行动,并随时间适应动态环境。这些挑战在现有的长时程智能体工作中并未被完全捕捉,因为后者通常考虑一个不随时间变化的静态环境。我们引入了ReLiveGym,一个用于长时程任务的诊断性评估环境,其中智能体在按时间顺序重放的数周真实世界新闻、市场和社交媒体流中稀疏地采取行动。这些任务涵盖了不同层次的时间敏感性、推理强度和重复性。在八种基础语言模型上,我们研究了模型选择与框架设计如何影响智能体在此类长时程任务上的表现。我们的结果表明,智能体如何决定何时行动成为长时程任务中一个重要的框架设计轴;且最优设计因任务而异,有时也因模型选择而异。我们还评估了从后见反馈中持续学习如何影响性能,并解决了在这些长时程任务中观察到的失败模式。这些发现表明,模型选择、行动时机机制和反馈利用是设计长时程智能体时的重要考量。代码:此https URL
英文摘要
As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic environment over time. These challenges are not fully captured in the existing long-horizon agent work, as they often consider a static environment that is not temporally changing. We introduce ReLiveGym, a diagnostic evaluation environment of long-lived tasks in which agents act sparsely over simulated weeks of chronologically replayed real-world news, market, and social-media streams. The tasks span diverse levels of time sensitivity, reasoning intensity, and recurrence. Across eight base language models, we investigate how model choice and harness design affect agent performance on such long-lived tasks. Our results show that how agents determine when to act arises as an important harness-design axis for long-lived tasks; and that the optimal design varies across tasks and sometimes model choices as well. We also evaluate how continuous learning from hindsight feedback affects performance and addresses failure modes observed in these long-lived tasks. These findings indicate model choice, action timing mechanism, and use of feedback as important considerations in the design of long-lived agents. Code: https://github.com/SaharaLabsAI/ReLiveGym
Comments9 pages. Preprint