发表机构
The Australian National University(澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究非标准无限时域动态规划中无限流评估的可行性,提出兼容性、递归压缩与广义求和框架,证明有限Hankel秩刻画时间非平稳性可被有限状态增广替代的条件,并对比不同求和准则的延续性。
AI 中文摘要
针对非标准无限时域准则的动态规划,不仅需要有限流值:一步控制替换必须保持可评估性,且延续信息必须递归传播。对于期望阶段回报流,我们区分了域可容许性、有限维连续自治可实现性、标量充分性以及上下文序一致性。对于绝对可和加权准则,我们证明了精确连续自治延续状态的最小欧几里得维数等于时间Hankel秩。因此,有限Hankel秩精确刻画了时间非平稳性何时能被有限内生状态增广所替代;下界甚至适用于非线性连续编码器,并且在一状态MDP中已然成立。对于字面意义的Candelpergher Ramanujan求和,相同的公共延续检验更早失效:仅改变当前动作就可能离开评估域。Abel求和提供了一个对比性的前缀封闭基准。
英文摘要
Dynamic programming for nonstandard infinite-horizon criteria requires more than finite stream values: one-step control substitutions must remain evaluable, and continuation information must propagate recursively. For expected stage-reward streams, we separate domain admissibility, finite-dimensional continuous autonomous realizability, scalar sufficiency, and contextual order consistency. For absolutely summable weighted criteria, we prove that the minimum Euclidean dimension of an exact continuous autonomous continuation state equals the temporal Hankel rank. Thus finite Hankel rank exactly characterizes when temporal nonstationarity can be replaced by finite endogenous state augmentation; the lower bound applies even to nonlinear continuous encoders and already holds in a one-state MDP. For literal Candelpergher Ramanujan summation, the same common-continuation test fails earlier: changing only the current action can leave the evaluation domain. Abel summation provides a contrasting prefix-closed benchmark.
Comments37 pages, including a 6-page electronic companion; 1 table; no figures