arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ContextWeave:一个真实世界工作流基准测试

ContextWeave: A Real-World Workflow Benchmark

Bo Wang, Yuqian Yao, Enxi Wang, Luozhijie Jin, Yang Liu, Yiran Suo, Yuxuan Cai, Enyu Zhou, Yufei Gao, Honglin Guo, Tianyu Huai, Li Ji, Zhikai Lei, Bufan Li, Lizhi Lin, Jinxiu Liu, Jie Yang, Jiazheng Zhou, Maosen Zhou, Pengfang Qian, Shichun Liu, Guanshan Liu, Hao Zheng, Yunhao Yu, Hang Yan, Jihua Kang, Xinchi Chen, Xipeng Qiu

arXiv 2608.04830首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute; ByteDance(复旦大学; 上海创新研究院; 字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出ContextWeave工作流基准测试,通过实验发现可操作的经验记忆能提升智能体工作流性能,为优化记忆系统提供了依据。

AI 中文摘要

当语言智能体从孤立任务转向长程、有状态的工作流时,记忆至关重要,但现有评估往往将其简化为检索或问答任务。我们推出ContextWeave,这是一个纵向基准测试,用于评估在现实办公工作流中,回忆的经验是否能提升下游智能体的性能。ContextWeave重构了14名参与者的隐私保护、历时数月的工作流,形成1005个可执行任务,其中包含568个核心评估任务,具备指令、容器化环境、轨迹以及任务特定的评分标准。它测量工作区质量以及与参与者特定偏好的一致性,同时补充了相关性、连续性、可解决性以及对误导性回忆的鲁棒性的诊断。在固定模型下的6个记忆组件中,最强配置将工作区评分从68.08提升至78.20,偏好评分从41.50提升至70.60。在固定记忆组件下,回忆能为所有5个测试的基础模型提升这两项结果,尽管提升幅度差异显著。我们的分析表明,可操作、经验丰富的记忆比紧凑摘要更能支持工作流延续并减少冗余探索,但也更容易受到误导性回忆的影响。这些发现为优化不仅是检索相关性,还有执行期间可靠使用的记忆系统提供了动力。

英文摘要

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑