Conservative Bias in Multi-Teacher Learning: Why Agents Prefer Low-Reward Advisors
多教师学习中的保守偏见:为什么智能体更偏好低奖励顾问
机构 * School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院) ; Escuela de Ingeniería, Universidad Central de Chile(智利中央大学工程学院)
专题命中 GUI与网页智能体 :agent(abstract);autonomous agent(abstract);分类 cs.AI
AI总结 本文揭示了多教师学习中智能体偏好保守低奖励顾问的现象,发现其受保守偏见主导,并在特定阈值下失效,同时在概念漂移情况下显著优于基线Q学习。
Comments 10 pages, 5 figures. Accepted at ACRA 2025 (Australasian Conference on Robotics and Automation)