发表机构
Human Inductive Bias Project(人类归纳偏差项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种基于强化学习智能体先验的无监督框架,用于从观察中检测环境中的智能体并重构其策略,解决了仅通过观察建模环境智能体的问题。
AI 中文摘要
我们研究人工智能系统如何仅通过观察识别并建模环境中的其他智能体,这是现实世界中协作行为所需的能力。该问题比逆强化学习约束更少,且在很大程度上未被探索。我们提出一个框架,使用在独立任务上训练的强化学习智能体作为智能体动态的先验,以执行智能体检测和策略重构。
英文摘要
We study how an AI system can identify and model other agents in its environment from observation alone, which is a capability necessary for cooperative behaviour in the real world. This problem is less constrained than inverse reinforcement learning and remains largely unexplored. We propose a framework that uses a reinforcement learning agent, trained on an independent task as a prior about agentic dynamics, to perform agency detection and policy reconstruction.