arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34419cs.AI

超越端到端黑箱映射:面向认知驱动面部反应生成的意图体框架

Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation

  • University of Exeter(埃克塞特大学)
  • William & Mary(威廉与玛丽学院)

机构由 AI 辅助整理,请以论文原文为准。

Hanzhong Zhang, Jindong Wang, Siyang Song

AI总结:

提出意图体框架,通过结构化内部状态和内在思维流将面部反应生成从端到端映射转为认知驱动生成,在REACT 2025上取得优异性能,并验证了自动评估指标的有效性。

AI中文摘要:

自动生成类人面部反应(FRG)对于构建能够进行人机交互(HCI)的智能系统至关重要。尽管多样且情境适当的面部反应能够反映人类互动中的潜在评价和情感过程,但现有大多数FRG方法依赖于端到端架构,直接将说话者行为映射到听者表情,而无需显式的中间内部状态。我们将FRG重新表述为由结构化内部状态过程介导的生成,并提出\ extbf{意图体}(Intentional Agent),它将FRG从直接的刺激-反应映射转变为通过显式中间状态进行的刺激基础生成。为了表示时间上的内部状态演化,我们提出了一种内部动力学模型,该模型将情感驱动力与迭代的内在思维流(ITF)整合到一个结构化的中间状态中,用于后续生成。该状态在对话沉默期间也能持续更新。此外,为了弥合抽象内部状态与生理动作之间的差距,我们将FRG表述为从潜在思维流到面部表情的下游情感映射。在REACT 2025数据集上的实验显示,FRDist为72.39,FRDiv为0.5057;感知合理性通过盲法人工评分单独评估。对96个反应的盲法人工评估发现,完整模型与真实反应的平均得分无显著差异(5.527对5.195,p_Holm=0.076),而完整模型显著优于事件触发模型和仅启发式模型(两者p_Holm<0.001)。反应质量评分器(RQS)与人工判断高度相关(Pearson r=0.855;Spearman ρ=0.821,两者p<0.05),支持其作为自动评估指标的使用。这些结果凸显了内生动力学在构建高度自主、类人智能体方面的巨大潜力。

英文摘要:

Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG methods rely on end-to-end architectures that directly map speaker behaviours to listener expressions without an explicit intermediate internal state. We reformulate FRG as generation mediated by a structured internal-state process and propose the \textbf{Intentional Agent}, which shifts FRG from direct stimulus-response mapping to stimulus-grounded generation through explicit intermediate states. To represent temporal internal-state evolution, we propose an internal dynamics model that integrates emotional drives with an iterative Inner Thought Flow (ITF) within a structured intermediate state used for subsequent generation. This state can continue to update during conversational silences. Furthermore, to bridge abstract internal states with physiological actions, we formulate FRG as a downstream affective mapping from this latent thought flow to facial expressions. Experiments on the REACT 2025 dataset show an FRDist of 72.39 and an FRDiv of 0.5057; perceptual plausibility is evaluated separately through blinded human ratings. A blinded human evaluation of 96 reactions found no significant difference in mean score between Full and ground truth ($5.527$ vs.\ $5.195$, $p_{\mathrm{Holm}}=.076$), while Full significantly outperformed Event-Triggered and Heuristic-Only (both $p_{\mathrm{Holm}}<.001$). The Reaction Quality Scorer (RQS) correlated strongly with human judgements (Pearson $r=.855$; Spearman $ρ=.821$, both $p<.05$), supporting its use as an automatic metric. These results underscore the immense potential of endogenous dynamics in building highly autonomous, human-like agents.

补充信息

↑