arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

即时场景图增长:应对长期机器人感知饱和

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics

Yue Chang, Rufeng Chen, Yifan Tian, Dazhi Huang, Zhaofan Zhang, Yi Chen, Wenze Zhang, Li Chen, Hui Xiong, Sihong Xie

arXiv 2607.13245首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Jilin University(香港科技大学(广州); 吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对长期机器人任务中传统3D场景图构建方式导致感知饱和的问题,提出JITOMA闭环框架,利用任务热图过滤观察、大语言模型解析意图动态唤醒锚点,经JITOMA-Bench评估,该框架有效减小图大小和延迟,保持处理时间稳定。

AI 中文摘要

虽然3D场景图为实体智能体提供了关键的结构化表示,但传统的提前构建所有内容然后过滤的管道与边缘平台的实时、低延迟需求相冲突,通过严重的观察冗余引发感知饱和效应。为解决此问题,我们提出了JITOMA(即时按需内存激活),这是一个将任务推理、感知和内存统一到即时增长过程中的闭环框架。JITOMA利用前端的自上而下任务热图过滤连续观察,路由最小流以维持低成本、休眠锚点的全局基础。在认知查询时,后端大语言模型解析机器人意图以动态唤醒与任务相关的锚点,仅在激活的本地子图内触发资源密集型操作。为评估这些动态能力并研究感知饱和权衡,我们引入了JITOMA-Bench,一个用于长期多任务和复杂多步推理的综合套件。广泛实验表明,JITOMA显著减小了活动图大小和字幕延迟,同时在长期任务切换下保持稳定的处理时间。

英文摘要

While 3D Scene Graphs (3DSGs) provide crucial structured representations for embodied agents, conventional Ahead-of-Time, "build-everything-then-filter" pipelines conflict with the real-time, low-latency demands of edge platforms, inducing a perceptual saturation effect via severe observation redundancy. To resolve this, we present JITOMA (Just-In-Time On-demand Memory Activation), a closed-loop framework that unifies task reasoning, perception, and memory into a just-in-time growth process. Instead of exhaustively mapping the entire environment, JITOMA leverages a top-down task heatmap at the frontend to filter continuous observations, routing minimal streams to maintain a global foundation of low-cost, dormant anchors. Upon a cognitive query, the backend Large Language Model (LLM) parses the robotic intent to dynamically awaken task-relevant anchors, triggering expensive semantic operations such as dense node captioning exclusively within the activated local subgraph. To evaluate these dynamic capabilities and study perceptual saturation trade offs, we introduce JITOMA-Bench, a benchmark for long-horizon task switching and complex intent grounding. Across JITOMA-Bench, JITOMA maintains only 1--6 active semantic objects and 0.25--0.28 s/frame across all tiers, showing that semantic computation remains bounded by current task demand rather than accumulated scene complexity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑