发表机构
The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体上下文过期导致的不一致问题,提出 Concord 框架,通过链接观察与来源、检测变化并策略化处理,在 ConcordBench 上实现与 oracle 一致的恢复效果,且节省 46.4% 令牌。
AI 中文摘要
现代智能体越来越多地依赖工具返回的观察结果(如从工作区读取的文件内容)来进行推理。然而,这些观察所依据的数据源之后可能被用户、其他智能体或外部工具修改,而模型仅在上下文中保留过时的内容。现有的智能体运行时几乎不支持通知模型先前观察到的某个事实已过期,导致智能体重用过时的观察结果,并对当前工作区状态做出错误断言。我们提出 Concord,一个上下文一致性框架,用于维护智能体上下文中的工具观察与其来源的可变数据源之间的一致性。Concord 将每个观察链接到其来源,检测来源变化,并使用可配置的处理策略在重用前更新、标注或抑制过时的上下文。Concord 适用于不同的智能体运行时和外部资源,并可轻松扩展到新的运行时-资源设置。我们将 Concord 实现为通用框架,并实例化一个具体用例来评估其有效性。我们构建了 ConcordBench,其中先前观察到的文件内容在后续编辑后变得过时。在三个评估的前沿模型中,Concord 在这些构造条件下所有评估案例中均产生了与恢复后工作区状态一致的答案,在该基准上恢复数量与 oracle 匹配,同时比最强的非 oracle 基线少使用 46.4% 的令牌。
英文摘要
Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agents, or external tools, while the model retains only the stale content in its context window. Existing agent runtimes provide little support for notifying the model that a previously observed fact has become stale, causing agents to reuse outdated observations and make incorrect claims about the current workspace state. We propose Concord, a context coherence framework that maintains the consistency between tool observation in agent context and the mutable sources from which they were derived. Concord links each observation to its source, detects source changes, and uses configurable handling policies to update, annotate, or suppress stale context before reuse. Concord is applicable across different agent runtimes and external resources, and can be easily extended to new runtime-resource settings. We implement Concord as a general framework, and instantiate a concrete use case to assess its effectiveness. We construct ConcordBench, where previously observed file contents become stale after subsequent edits. Across three evaluated frontier models, Concord produces answers consistent with the restored workspace state in all evaluated cases under these constructed conditions, matching the oracle on recover count for this benchmark, while using 46.4% fewer tokens than the strongest non-oracle baseline.
Comments8 pages, 3 figures, 1 table. Accepted to the AgenticOS Workshop at SOSP 2026