arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人类行为中的神经符号层次意图预判

Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior

Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat, Ghazaleh Khodabandelou

arXiv 2609.17064首次发表:更新:

发表机构

LISSI Laboratory; University of Paris-Est Créteil (UPEC); IUT de Créteil-Vitry(LISSI实验室; 巴黎东部克雷泰伊大学(UPEC); 克雷泰伊-维特里大学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出神经符号层次意图预判方法,通过层次规划解码器结合软正则化与硬掩码,在四层基准上实现优于基线的预判,并保证逻辑有效性。

AI 中文摘要

辅助自主系统必须在观察到行为完成之前预判人类目标。本文将预判问题形式化为从部分观察的多模态片段中进行目标推断,并结合对剩余行为的结构化预测,而非精确的运动预测。一个紧凑的层次规划解码器(HPD)被附加到一个冻结的神经符号识别编码器上,并在四个本体层级上预测下一步动作、剩余活动和低层意图,以及片段的高层意图(HLI)。该解码器通过结合转移一致性和层次连续性损失的软神经符号正则化进行训练,并在推理时使用硬可达性掩码来强制本体有效性。在一个基于NTU RGB+D 120特征构建的包含15,002个多模态片段的组合式四层基准上,三个显著特性同时被观察到。相对于最强的序列基线,优势随预判时域增加而扩大,从第1步的+1.7个百分点到第3步的+7.3个百分点(top-5)。在组合泛化设置下,即每个多父低层意图保留一个父关联时,该优势在第1步扩大至+4.9个百分点。在片段层级,96.8%的预判轨迹满足联合逻辑约束,高于最强基线的88.1%和真实标签下限的73.9%;仅软逻辑项就使HLI可达性违规相对减少59.8%至71.1%,而硬掩码则将其完全消除。神经生成提供预测排序,符号约束提供逻辑有效性,二者的结合产生了连贯的层次预判,同时暴露了组合目标泛化和无序集合预测方面的剩余挑战。

英文摘要

Assistive autonomous systems must anticipate human goals before an observed behavior is complete. This article formulates anticipation as goal inference from a partially observed multimodal episode together with structured prediction of the remaining behavior, rather than exact motor forecasting. A compact Hierarchical Planning Decoder (HPD) is attached to a frozen neuro-symbolic recognition encoder and predicts, at four ontological levels, the next actions, the remaining activities and low-level intentions, and the episode high-level intention(HLI). The decoder is trained with soft neuro-symbolic regularization combining transition-coherence and hierarchical continuity losses, and is decoded with hard reachability masks that enforce ontological validity at inference. On a compositional four-level benchmark of 15,002 multimodal episodes built over NTU RGB+D 120 features, three headline properties are observed together. The advantage over the strongest sequential baseline grows with the anticipation horizon, from +1.7 points at step 1 to +7.3 points at step 3 (top-5). Under compositional generalization, where one parent association per multi-parent low level intention is held out, this advantage widens to +4.9 points at step 1. At the episode level, 96.8% of anticipated trajectories satisfy the joint logic constraints, above the 88.1% strongest-baseline value and the 73.9% ground-truth floor; soft logic terms alone account for a 59.8 to 71.1% relative reduction of HLI-reachability violations, and the hard masks then eliminate them entirely. Neural generation supplies predictive ranking, symbolic constraints supply onto logical validity, and their combination yields coherent hierarchical anticipation while exposing remaining challenges in compositional goal generalization and unordered set prediction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑