arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RECON:用于长上下文组合推理的智能体内存基准测试

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

Mihir Shriniwas Arya

arXiv 2607.16716首次发表:更新:

发表机构

RV College of Engineering; Department of Computer Science and Engineering(RV工程学院; 计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究引入RECON基准测试,用于评估长上下文组合推理,跨越三个领域24个案例文件,测试智能体六个内存密集型任务,揭示当前架构在内存相关任务上存在重大局限,非预言机系统准确率低,检索和推理有挑战。

AI 中文摘要

大语言模型和基于大语言模型的智能体被广泛用作个人聊天助手、企业助手和自主工作流智能体。在所有这些应用中,内存(即在长上下文和多次交互中保留、访问和推理积累信息的能力)在决定智能体的可靠性方面起着关键作用。我们引入了RECON(带有模糊叙述的扩展上下文推理),这是一个用于评估长上下文组合推理的基准测试。RECON跨越三个领域(刑事、医疗和金融)的24个案例文件,每个文件从50k到100k令牌不等,并在六个内存密集型任务上测试智能体:重建多跳证据链、传播级联无效、解决源冲突、反事实推理、满足时间约束和时间事实检索。最近的内存基准测试评估智能体是否能检索分散的事实或检测事实是否发生变化,而RECON评估变化后会发生什么,智能体是否能追踪哪些下游结论受到影响,哪些通过独立支持得以保留,以及替代时间线会如何展开。我们的评估揭示了当前架构存在的重大局限性:即使是最强的非预言机系统准确率也仅达到22.4%,检索和推理都面临挑战。

英文摘要

Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents. In all these applications, memory (the ability to retain, access, and reason over information accumulated over long contexts and multiple interactions) plays a crucial role in determining the reliability of any agent. We introduce RECON (Reasoning over Extended Contexts with Obfuscated Narratives), a benchmark for evaluating compositional reasoning over long contexts. RECON spans 24 case files across three domains (criminal, medical, and financial), each ranging from 50k to 100k tokens, and tests agents on six memory intensive tasks: reconstructing multi-hop evidence chains, propagating cascading invalidations, resolving source conflicts, counterfactual reasoning, satisfying temporal constraints, and temporal fact retrieval. Recent memory benchmarks evaluate whether agents can retrieve scattered facts or detect if a fact has changed whereas RECON evaluates what happens after the change, whether agents can trace which downstream conclusions are affected, which survive through independent support, and how alternative timelines would have unfolded. Our evaluation reveals substantial limitations across current architectures: even the strongest non-Oracle system reaches only 22.4% Accuracy, with retrieval and reasoning each surfacing as challenges.

Comments25 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑