The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
霍克鲁斯:用于检测和缓解具身AI系统中奖励黑客行为的可解释任务分解
机构 * Berkeley AI Safety Initiative (BASIS) UC Berkeley(伯克利人工智能安全计划(BASIS)伯克利大学) ; Department of Electrical and Computer Engineering Johns Hopkins University(电气与计算机工程系约翰霍普金斯大学)
专题命中 模仿学习与强化学习 :embodied AI(title,abstract);分类 cs.LG;robotic(comments)
AI总结 本研究提出MITD方法,通过可解释性任务分解有效检测和缓解具身AI系统中的奖励黑客行为,实验表明分解深度可显著降低奖励黑客频率。
Comments Accepted to the NeurIPS (Mexico City) 2025 Workshop on Embodied and Safe-Assured Robotic Systems (E-SARS). Thanks to Aman Chadha