表示充分性的自认证:最小任务损失下的序列认证
Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss
AI总结:
本文针对智能体历史压缩表示的混叠风险,提出四层自认证理论,构建序列认证策略并证明其成本渐近匹配下界,为表示充分性的认证提供了理论框架。
AI中文摘要:
基于历史压缩表示行动的智能体面临结构性风险:若该表示将具有不同最优行动的历史混叠,则任何仅基于该表示可测量的规则都无法避免不可约的每轮损失,且智能体可能无法从自身记录中检测到这一点。本文提出了表示充分性自认证的四层理论:静态层通过贝叶斯风险分组恒等式定义决策论充分性,并通过精确总变差阈值定价一次性外部验证;序列层将认证建模为任务损失货币下的最优停止问题,通过覆盖线性规划定义环境相关的认证复杂度常数,证明了每个δ-正确策略的信息-任务损失下界,并给出了认证跟踪停止策略,其成本渐近匹配该下界;最终边界层给出了显式核切换示例,并指出了覆盖策略切换或表示修复所需的开放定理,未声称固定核保证可扩展至表示修正。两个主要定理的完整证明见附录。
英文摘要:
Agents that act on a compressed representation of their history face a structural risk: if the representation aliases histories with different optimal actions, no rule measurable with respect to the representation can avoid an irreducible per-round loss, and the agent may be unable to detect this from its own transcript. This paper develops a four-layer theory of self-certification of representation adequacy. The static layer defines decision-theoretic adequacy through a Bayes-risk grouping identity and prices a one-shot external verification by an exact total-variation threshold. The sequential layer poses certification as an optimal-stopping problem in the currency of task loss: we define an environment-wise certification complexity constant through a covering linear program, prove an information-task-loss lower bound for every delta-correct strategy, and give a Certification Track-and-Stop policy whose cost matches the bound asymptotically. A final boundary layer gives an explicit kernel-switching example and identifies the open theorem needed to cover policy switching or representation repair; it does not claim that the fixed-kernel guarantees extend to representation revision. The proofs of the two main theorems are given in full in the appendices.