改变行动的記憶并非指导行动的記憶:历史条件化机器人策略的反事实审计
Memory That Changes Action Is Not Memory That Guides It: Counterfactual Auditing of History-Conditioned Robot Policies
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- ShanghaiTech University(上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出反事实记忆审计(CMA),通过交叉历史评估机器人策略,证明记忆改变行动不等于指导行动,并区分记忆敏感性与决策可靠性。
AI中文摘要:
一个将积木放回原位的机器人可能会遇到两个任务一致的历史,它们汇聚到相同的当前输入,但需要不同的行动。然而,记忆策略的评估往往依赖于任务成功或记忆扰动下的行动变化,这两者都不能证明记忆指导了决策。我们提出了反事实记忆审计(CMA),一种评估协议,它在验证一致的当下交叉两个历史,在共同随机性下查询冻结策略,并在两种历史下评估每个保存的行动。这区分了记忆敏感性、合理选择、匹配世界的物理价值和每对可靠性。在Mem-0上,每个被审计的放回对都改变了行动,但只有20/64对完全可靠;在稍后的交换决策中,所有配对行动都改变,而两种记忆选择了相同的分支。原生干预进一步显示了闭环影响:替换历史库将行为重定向到替换内容,而恢复一个4096字节的保护锚点恢复了因注入库故障而丢失的38.9点交换成功率。在双臂物理平台上,记忆改变了保存的行动,但九个完成的放回操作中有五个达到了错误的目标。这些结果表明,机器人可以记住并反应,而不必可靠地使用记忆来选择其过去所保证的行为。CMA提供了一种决策级审计来区分这些情况。
英文摘要:
A robot returning a block to its origin tray may encounter two task-consistent pasts that reconverge to the same current input but warrant different actions. Yet memory-policy evaluations often rely on task success or action change under memory perturbation, neither of which establishes that memory guides the decision. We propose the \textbf{Counterfactual Memory Audit (CMA)}, an evaluation protocol that crosses two histories at a verified-identical present, queries a frozen policy under common randomness, and evaluates each saved action under both pasts. This separates memory sensitivity, warranted choice, matched-world physical value, and per-pair reliability. On Mem-0, every audited Put Back pair changes action, but only $20/64$ pairs are fully reliable; at a later Swap decision, all paired actions change while both memories select the same branch. Native interventions further show closed-loop influence: replacing the history bank redirects behavior toward the replaced content, while restoring a 4096-byte protected anchor recovers $38.9$ points of Swap success lost to injected bank faults. On a dual-arm physical platform, memory changes saved actions, yet five of nine completed Put Back manipulations reach the wrong target. These results show that a robot can remember and react without reliably using memory to choose the behavior its past warrants. CMA provides a decision-level audit for distinguishing these cases.