关于强化学习开发智能体自定义行为的注记
A Note on Reinforcement Learning to Develop Self-defined Agents' Behavior
浏览论文内容
中文总结 AI 辅助
本研究基于有限理性假设,采用无监督学习为主的技术,通过Cross Targets(交叉目标)方法训练人工智能体,实现行为内部一致性,成功让随机游走智能体解决图节点分类问题。
中文摘要 AI 辅助
本注记的核心是在有限理性假设下开发简单行为策略的自适应性。我们的人工智能体采用以无监督学习为主的学习技术,以实现行为的内部一致性,取得了意外结果。这些结果主要可视为观察者解释的效应。采用的第一种技术名为Cross Targets(交叉目标):为训练学习智能体,我们使用关于待执行动作的预测与关于后续结果的预测之间交叉的数据。随后呈现了CT(交叉目标)“盲”策略开发的应用:随机游走智能体在学习如何停留在同质区域后,解决了图上的节点分类问题。
英文摘要
The key point in this note is the self-development of simple behavior strategies, consistently with the bounded rationality hypothesis. Our artificial agents adopt learning techniques, mainly unsupervised, to achieve internal consistency in their behavior, with unexpected results. Those results can be considered mainly as the effects of the observer interpretation. The first technique in use has the name Cross Targets: to train the learning agent we use data crossed between the guesses about the action to be done and the guesses about the following results. An application of the CT "blind" strategy development is then presented: random walkers solve a node classification problem on a graph, after having learnt how to remain in homogeneous regions.