arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

情感偏好作为目标优先级调节

Emotional Preferences as Goal-Priority Regulation

Shiqi Liu, Yihua Tan, Hu Fu, Guanyu Qi

arXiv 2608.27072首次发表:更新:

发表机构

School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出情感偏好作为目标优先级调节的概念,构建含多目标强化学习内部控制器与外部偏好生成器的框架,经实验验证其偏好函数表现优于固定偏好与手工设计偏好策略。

AI 中文摘要

智能体决策的核心问题在于,竞争的低层目标的相对优先级是否可由高层目标自主生成的情感偏好确定,而非外部预先指定。在变化的外部环境与演化的内部状态下,情感在调节竞争目标的相对优先级方面发挥重要功能作用。受情感目标导向理论启发,本文研究此类偏好调节如何通过强化学习在计算上实现。我们首先提出一种涌现情感偏好的概念:高层目标自主诱导竞争低层目标的、依赖状态的偏好。该概念基于由多目标强化学习内部控制器与外部偏好生成器组成的框架构建。内部控制器提供一系列依赖偏好的目标导向行为,外部偏好生成器通过对高层目标的强化学习,学习从当前状态到目标偏好的映射。我们将情感偏好操作化为通过优化涌现的、依赖状态的目标相对优先级调节。此外,我们刻画了由偏好调节诱导的策略空间,并根据内部行为库的表示误差推导最优性差距的上界。当最优策略可由可用的依赖偏好策略表示时,该差距消失。在自建的多目标探索环境中的实验表明,所学偏好函数表现出情境优先级切换、分级权衡与时间持续性,且优于评估的固定偏好与手工设计偏好策略。

英文摘要

A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. Under changing external environments and evolving internal states, emotions play an important functional role in regulating the relative priorities of competing goals. Inspired by the goal-directed theory of emotion, this paper studies how such preference regulation can be computationally realized through reinforcement learning. We first propose a conception of emergent emotional preference: a high-level goal autonomously induces state-dependent preferences over competing lower-level objectives. This conception is built upon a framework consisting of a multi-objective reinforcement learning inner controller and an outer preference generator. The inner controller provides a repertoire of preference-conditioned goal-directed behaviors, while the outer preference generator learns a mapping from the current state to objective preferences through reinforcement learning on a high-level goal. We operationalize emotional preference as a state-dependent regulation of relative goal priorities that emerges through optimization. Furthermore, we characterize the policy space induced by preference regulation and derive an upper bound on the optimality gap in terms of the representation error of the inner behavioral repertoire. We show that the gap vanishes when the optimal policy can be represented by the available preference-conditioned policies. Experiments in self-constructed multi-objective exploration environments show that the learned preference function exhibits contextual priority switching, graded trade-offs, and temporal persistence, and outperforms the evaluated fixed-preference and handcrafted-preference strategies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑