arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07420cs.CV

统一视觉中心的行人过街动作预测:自适应补丁投影与主动空间校正

Unified Vision-Centric Pedestrian Crossing Action Prediction via Adaptive Patch Projection and Proactive Spatial Rectification

  • School of Electronics and Control Engineering, Chang’an University(长安大学电子与控制工程学院)
  • School of Astronautics, Northwestern Polytechnical University(西北工业大学航天学院)

机构由 AI 辅助整理,请以论文原文为准。

Yao Tian, Le Yang, Binglu Wang

AI总结:

ViCross提出一种基于多模态大语言模型的视觉中心行人过街动作预测框架,通过可变分辨率补丁映射和空间约束增强策略,在无需额外感知模块下实现高效准确的预测,性能优于现有方法。

AI中文摘要:

视觉线索对于行人动作预测是可获取且信息丰富的,但在没有帧级外部感知线索的情况下,从视频帧中获取稳定的以目标为中心的表示仍然具有挑战性。因此,大多数方法依赖于额外的感知模块或多源信息融合,使得视觉中心设置的可信度成为一个未解决的问题。为此,我们提出了ViCross,一种由多模态大语言模型驱动的视觉中心行人过街动作预测框架,该框架在除首帧目标初始化外无需额外感知模块的情况下,从视频帧中维持以目标为中心的推理。尽管多模态大语言模型展现出强大的视觉理解能力,但将其直接应用于视觉中心动作预测面临两个挑战。首先,准确感知目标行人通常需要高分辨率输入和密集的视觉令牌化,这使得全帧编码在计算上不可行。ViCross通过可变分辨率补丁映射模块来解决这一问题,实现高效的令牌分配,同时保留关键的行人细节。其次,缺失的时空先验阻碍了跨帧一致推理。ViCross通过空间约束增强策略来缓解这一问题,该策略捕获过去的运动、未来的位置和动作语义,用于训练时的主动空间校正。大量实验表明,ViCross在视觉中心预测设置中带来了明显的性能提升,并在多种设置下与多源融合方法具有竞争力。代码可在以下网址获取:此https URL。

英文摘要:

Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centric representations from video frames remains challenging without frame-level external perception cues. Thus, most methods rely on additional perception modules or multi-source information fusion, leaving the reliability of vision-centric setting an open question. To this end, we propose ViCross, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization. While multimodal large language models exhibit strong visual understanding, applying them directly to vision-centric action prediction faces two challenges. First, accurately perceiving target pedestrians often requires high resolution inputs and dense visual tokenization, making full-frame encoding computationally prohibitive. ViCross tackles this with Variable Resolution Patch Mapping module for efficient token allocation while preserving key pedestrian details. Second, missing spatiotemporal priors hinder consistent cross frame reasoning. ViCross mitigates this with a Spatial Constraint Enhancement Strategy that captures past motion, future locations, and action semantics for training-time proactive spatial rectification. Extensive experiments show that ViCross delivers clear gains in vision-centric prediction settings and is competitive with multi-source fusion approaches in several settings. Code is available at https://github.com/2tianyao1/ViCross.git.

↑