发表机构
School of Health & Wellbeing, University of Glasgow; Department of Respiratory and Critical Care Medicine, Shanghai Sixth People’s Hospital, Shanghai Jiao Tong University School of Medicine; School of Life Science and Technology, University of Electronic Science and Technology of China; Institute of Health Informatics, University College London(格拉斯哥大学健康与福祉学院; 上海交通大学医学院附属第六人民医院呼吸与危重症医学科; 电子科技大学生命科学与技术学院; 伦敦大学学院健康信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能系统从自我修订记录检索信息时面临的证据状态修订问题,通过比较多种方法,发现渲染对评估有混淆影响,提出内存评估应固定渲染,弃用感知系统应采用覆盖查询的最粗保留状态。
AI 中文摘要
人工智能系统越来越多地从自我修订的记录中检索信息,如问题线程、百科历史、政策日志和长对话等。挑战不仅在于找到相关证据,还在于确定哪些声明仍然有效、哪些已被取代以及何时弃权。结构化内存有望通过类型化边、时间更新和冲突状态来解决此问题,但评估通常会同时改变机制和提示呈现。我们将此作为证据状态修订进行研究,在来自GitHub、多仓库问题历史、维基百科和DyKnow风格时间流的2907个高度一致的问题上比较了平面检索、粗粒度边失效和细粒度修订账本。一个渲染匹配的对照组(相同布局,禁用弃用)揭示了核心混淆:当一个值被更改并随后恢复时,修订账本似乎比平面基线高出0.182,但几乎所有增益都来自更简单的呈现;细粒度机制残余与零无显著差异(两个评判组中为+0.021至+...
英文摘要
AI systems increasingly retrieve from records that revise themselves: issue threads, encyclopedic histories, policy logs, and long conversations. The challenge is not only finding relevant evidence, but deciding which claims remain in force, which were superseded, and when to abstain. Structured memories promise to solve this with typed edges, temporal updates, and conflict status, yet evaluations often change mechanism and prompt presentation together. We study this as Evidence-State Revision, comparing flat retrieval, coarse edge invalidation, and fine-grained RevisionLedger on 2,907 high-agreement questions from GitHub, multi-repo issue histories, Wikipedia, and DyKnow-style temporal streams. A render-matched control (same layout, deprecation disabled) reveals the central confound: when a value is changed and later restored, RevisionLedger appears to beat a flat baseline by +0.182, but almost all the gain comes from easier presentation; the fine-grained mechanism residual is indistinguishable from zero (+0.021 to +0.025 across two judge families). After presentation is controlled, coarse invalidation is the only mechanism that pays for current-state queries, beating the fine ledger by 0.084; the same query-sufficiency principle says provenance mainly needs retained invalidated evidence, not richer typing. Memory evaluations should hold render fixed, and deprecation-aware systems should deploy the coarsest retained state that covers their queries.