arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12486eess.SYcs.AIcs.GTcs.SYmath.OC

自适应智能体设计

Adaptive Agent Design

Raj Kiriti Velicheti, Subhonmesh Bose, Tamer Başar

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究智能体在非马尔可夫环境中的双层设计问题,提出通过软Q学习优化转移核与策略,并分析其收敛性及与最优策略的差距。

中文摘要 AI 辅助

我们考虑一个智能体在一般的非马尔可夫环境中采取行动。该智能体维护其自身的智能体状态,但可以自由选择这些状态之间的转移核,并优化其状态反馈控制策略。我们研究双层智能体设计问题,该问题在给定转移核的情况下,利用通过行为策略获得的观测和动作的离线数据,优化转移核及其所诱导的策略。对于一般环境,我们证明软Q学习算法几乎必然收敛到由行为策略和所选转移核诱导的平稳平均值所定义的软贝尔曼方程的固定点,并阐明所得策略与最优策略之间的区别。在部分可观测马尔可夫决策问题中,我们通过零阶和贝叶斯优化技术分析参数化转移核设计的收敛性质。

英文摘要

We consider an agent acting against a general non-Markovian environment. The agent maintains its agent states, but is free to choose a transition kernel across those states and optimize its state-feedback control policies. We study the bi-level agent design problem that optimizes the transition kernel and the policy it induces, given said kernel with offline data of observations and actions obtained via a behavioral policy. For general environments, we show that a soft $Q$-learning algorithm converges almost surely to the fixed point of a soft Bellman equation defined by the stationary averages that the behavioral policy and the chosen kernel induce, and we delineate what separates the resulting policy from an optimal one. In partially observed Markov decision problems, we analyze convergence properties of parametrized transition kernel design via zero-th order and Bayesian optimization techniques.

发表机构

  • University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑