发表机构
Zhejiang University; Ant Group(浙江大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OneDayAgent是一种长程管控框架,可将开放式请求分解为子任务,在不同后端LLMs上实现最优性能,解决智能体的多类失效模式。
AI 中文摘要
大语言模型(LLM)智能体正越来越多地被应用于覆盖工作、学习和生活的开放式日常请求中。这些任务具有长程性、跨环境性和多模态性,要求智能体在经过异构工具和附件时,能在多步骤中保留目标与约束。尽管已有研究解决了目标漂移、状态丢失、上下文溢出等个别失效模式,但针对单一管控框架能否同时管理这些问题并在不同后端保持有效性的研究较少。我们提出OneDayAgent,这是一种面向自主智能体的长程管控框架。OneDayAgent将开放式请求转化为受管控的执行流程,该流程将任务分解为有边界的子任务,在上下文压力下维护执行记忆,并验证和修复最终交付物。我们在AgentIF-OneDay数据集上对OneDayAgent进行了104项任务的评估。使用GLM-5.2后端时,OneDayAgent取得了0.821的总体得分,创下了新的最优水平。该管控框架可在来自三个模型家族的五个后端大语言模型上运行,表明其无需微调即可跨后端泛化,即使不同模型在同一工作流下会产生不同的执行风格。
英文摘要
LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous tools and attachments. While prior work has addressed individual failure modes such as goals drift, states loss, and context overflow, whether a single harness can manage them jointly and remain effective across backends has received less study. We present OneDayAgent, a long-horizon harness for autonomous agents. OneDayAgent turns an open-ended request into a managed execution process that decomposes tasks into bounded subtasks, maintains execution memory under context pressure, and verifies and repairs the final deliverable. We evaluate OneDayAgent on AgentIF-OneDay across 104 tasks. With the GLM-5.2 backend, OneDayAgent sets a new state of the art with an overall score of 0.821. The same harness runs across five backend LLMs from three model families, indicating the harness generalizes across backends without tuning, even as different models induce distinct execution styles under the same workflow.
CommentsOngoing work