用于内容推荐的反向心智理论建模:从网页浏览到动态智能界面
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
- JPMorganChase(摩根大通)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出反向心智理论(IToM)流程,从用户交互反向推理推断信念与偏好,在OPeRA数据集上验证其画像推断效果,并通过VisionOS应用展示跨模态迁移能力,提升内容推荐的用户理解精度。
AI中文摘要:
现代推荐系统将观测到的用户行为视为用户偏好的可靠替代指标,但用户交互往往反映的是探索或比较行为,而非稳定的偏好表达。随着界面从静态布局向生成式用户界面(UI)和沉浸式扩展现实(XR)演进,对更深入、模态无关的用户理解的需求日益增长:这些自适应环境不仅需要决定呈现什么内容,还需决定在何处、何时、以何种突出程度呈现,以及最重要的——用户采取行动的原因。我们提出一种反向心智理论(IToM)流程,该流程从观测到的交互反向推理,以推断能解释用户行为的信念、偏好和决策特质。该流程重构每个用户的决策情境,包括被选中的内容以及可用的替代选项,应用大语言模型(LLM)驱动的反事实推理生成基于证据的自然语言信念陈述,并通过多假设溯因推理将这些信念合成为结构化用户画像。我们在OPeRA数据集上,针对四个任务(下一个动作预测、购物态度对齐、大五人格推断以及保留类别预测),与真实人格评估、态度调查和基于访谈的画像进行对比评估。结果显示,推断出的画像与真实画像匹配度更高或表现更优,且多假设推理对准确的人格预测至关重要。我们还通过VisionOS上的一个由画像驱动的空间银行应用,展示了跨模态迁移能力。
英文摘要:
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.