arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22847cs.CV

注视锚定社交网络:通过联合建模解码隐式关系

Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

  • University of Birmingham(伯明翰大学)
  • Korea Electronics Technology Institute(韩国电子技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yuqi Hou, Zhuo Chen, Han Hu, Je Woo Kim, Jianbo Jiao, Hyung Jin Chang

AI总结:

研究针对标准模型处理注视问题的不足,提出ANCHOR范式,通过联合建模视觉注意力和隐式关系解码注视锚定的社会意图,利用关系注意力机制等方法,在扩展基准上验证,取得先进性能,证明可从静态注视模式学习隐式社会层次结构。

AI中文摘要:

人类注视不仅指向视觉目标,在静态图像中它还是社会意图的微妙指标,而标准模型通常独立处理个体,将注视视为独立同分布量或孤立预测社会语义。近期多人方法存在不足。我们提出ANCHOR,一种以目标为中心的范式,通过对视觉注意力和潜在隐式关系的联合分布建模来解码注视锚定的社会意图。该方法利用关系注意力机制捕捉细粒度人际联系,通过特征调制进行多人解析。为稳定训练,实现优化协同以解决空间注视精度和潜在社会推理间的冲突。在扩展基准上验证,结果显示了先进性能,首次定量证明可从静态注视模式中稳健地解缠并学习隐式社会层次结构。

英文摘要:

Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals independently, treating gaze as an i.i.d. quantity or predicting social semantics in isolation. Recent multi-person methods attempt to address this but often treat social relations as rigid, post-hoc classifications decoupled from the gaze estimation process. This oversimplification fails to capture the nuanced nature of social intent, which acts as an underlying driver of gaze behavior rather than a secondary categorical output. We address these limitations by proposing ANCHOR, a target-centric paradigm designed to decode gaze-anchored social intent by modeling the joint distribution of visual attention and latent implicit relations. Our approach surfaces these dependencies as the latent structural scaffolding of gaze behavior. The architecture utilizes a relational attention mechanism to capture fine-grained interpersonal links, leveraging feature-wise modulation for efficient multi-person parsing from a single vision backbone. To stabilize the training of this coupled formulation, we implement an optimization synergy to resolve the inherent conflicts between spatial gaze accuracy and latent social reasoning. This approach ensures robust generalization by seeking stable, flat minima while simultaneously harmonizing competing task gradients. We validate our framework on an extended benchmark featuring dense multi-person annotations and novel social influence rankings. Our results demonstrate state-of-the-art performance and provide the first quantitative evidence that implicit social hierarchies can be robustly disentangled and learned directly from static gaze patterns.

补充信息

↑