EgoSIS:从分解的视觉自我转换到运动规范空间证据用于无人机推理
EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
浏览论文内容
中文总结 AI 辅助
EgoSIS通过分解视觉自我转换和运动规范空间证据,为无人机视频问答提供无需位姿的适配器,在SIS-Bench上显著提升感知与记忆准确率。
中文摘要 AI 辅助
无人机视频问答需要将相机运动与场景变化分离,但仅依赖RGB的多模态模型缺乏明确且稳定的参考来实现这种分离。我们提出EgoSIS,一种无需位姿的适配器,通过三个阶段将RGB导出的双向光流转换为运动规范视觉证据。分解的视觉自我转换(FVET)拟合鲁棒的图像平面转换,并暴露运动、残差支持和可靠性因子。可靠性门控的自我转换记忆(ReTEM)使用可靠性加权更新来维护有界历史,并在剪辑或持续不确定性时重新锚定。自我对齐空间证据(EASE)将支持的视觉特征扭曲到每个片段的局部锚点,并通过零初始化残差在每个视觉切片中注入四个空间证据令牌,而不改变Qwen的视觉令牌数量。在SIS-Bench上,EgoSIS-8B获得89.9%的感知准确率、82.5%的感知加记忆准确率和76.2%的整体准确率,最大增益集中在自我意识感知和记忆方面。该适配器因此提供了光流与空间推理之间的可解释接口。
英文摘要
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. Factorized Visual Ego-Transitions (FVET) fits a robust image-plane transition and exposes motion, residual-support, and reliability factors. Reliability-Gated Ego-Transition Memory (ReTEM) uses reliability-weighted updates for a bounded history and re-anchors it at cuts or sustained uncertainty. Ego-Aligned Spatial Evidence (EASE) warps supported visual features into each segment's local anchor and injects four spatial evidence tokens per visual slice through zero-initialized residuals, without changing Qwen's visual-token count. On SIS-Bench, EgoSIS-8B obtains 89.9\% perception, 82.5\% perception-plus-memory, and 76.2\% overall accuracy, with the largest gains concentrated in self-awareness perception and memory. The adapter thus provides an interpretable interface between optical flow and spatial reasoning.
发表机构
- Beihang University(北京航空航天大学)
- Zhongguancun Academy(中关村学院)
- Northeastern University(东北大学)
- Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间应用工程与技术中心)
机构由 AI 辅助整理,请以论文原文为准。