发表机构
Indian Institute of Technology Ropar(印度理工学院罗巴尔校区)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Nous框架,证明在隐马尔可夫模型中,学习并认证智能体记忆决策所需的记录数可比源校准少二次方量级,并通过MiniGrid实验验证了无需源校准即可可靠认证决策改进。
AI 中文摘要
基于信念的智能体记忆需要对当前状态做出可靠决策,但其证据可能是噪声的、复制的或过时的。记忆是否必须先对其来源进行校准,才能改进其决策?我们将学习、校准和修订认证分离开来。在一个四模型隐马尔可夫族上,学习一个未知的贝叶斯决策需要Theta(l^-2)条记录,而认证其相对于一个有信息的现任者的改进需要来自同一观测规律的O(l^-2)条新记录,而固定精度的源估计则需要Theta(l^-4)条记录,当持久性l趋于零时。因此,学习和认证有用的决策所需的记录数量可能比源校准少二次方量级。一个更广泛的模型类保留了决策速率和源下界。在未知的同一性加背景报告信道下,我们刻画了策略改进的尖锐识别区间,并利用可观测的见证区域推导出有限样本证书,无需纯类锚点。一个鲁棒性扩展容忍有界的历史依赖误设定和条件复制;分裂训练的见证适用于任意历史空间,并具有明确的功效条件。我们将策略边界收据与Nous Dimensions集成,并在三个外部MiniGrid记忆环境中测试了45,000个保留的可变状态历史和9,000个回合,引入了噪声报告接口。新证书接受9/9个相对于常数现任者的改进和4/9个相对于最后写入获胜的改进,而早期证书在MiniGrid中则没有接受任何改进。强大的既有推断基线仍具有竞争力或更优。结果是关于何时可以在不恢复源可靠性的情况下学习和证明记忆决策的统计说明,而非一种普遍优越的记忆算法。
英文摘要
Agent memory systems update state decisions from reports whose reliability may be unknown. Existing analyses of source estimation do not determine when a policy can be learned or its improvement certified without identifying the reporting channel. We study these three tasks using the same observed records. For a specified hidden Markov family with continuous source uncertainty, decision learning and powered certification have quadratic sample complexity, whereas fixed-precision source estimation has quartic complexity. We characterize a sharp identified interval for policy gain under an unknown shared-background channel and derive finite-sample certificates under bounded history dependence and conditional copying. Independently trained witness regions support general history spaces, and disagreement-conditioned auditing improves power for sparse revisions. For dependent histories, prediction-count-preserving batches cancel the unknown reporting background and admit conditional certificates. A MultiWOZ 2.4 evaluation uses text-processing policies on 1,000 human-written test dialogues with simulated audits. Balanced batches retain 2.26 percentage points of the full candidate's 7.34 percentage-point mean gain and obtain more positive certificates under weak audits. These results establish task-specific information requirements and provide an auditable policy-revision framework for Nous.
Comments43 pages, 7 figures, including appendices. Revised and extended theory and evaluation, with prediction-count-preserving batch certificates and a MultiWOZ 2.4 text-policy study using simulated audits. Code and reproducibility archive: https://github.com/Pranavsingh431/nous-state/releases/tag/jmlr-submission-2026-10-04