arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoGazeLite:面向 token 高效多模态大语言模型视频输入的设备内自我中心注视预测

EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input

Matteo Stoiber, Niels Buus Lassen

arXiv 2608.15614首次发表:更新:

发表机构

Copenhagen Business School(哥本哈根商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EgoGazeLite 是轻量双进程注视预测器,可在消费级硬件上实时运行,无需眼动追踪硬件即可实现 token 高效的自我中心视频理解,性能与真实注视裁剪无显著差异。

AI 中文摘要

多模态大语言模型(MLLM)用于可穿戴设备的自我中心视频理解时,受 token 预算限制:内存和计算成本随视觉 token 数量增加而上升,高分辨率视频的大规模传输与处理成本高昂。现有研究 GazeLLM 通过围绕佩戴者注视裁剪视频,将视觉 token 数量减少约 10 倍,同时保持或提升全分辨率描述质量,但该压缩策略依赖专用眼动追踪硬件,而消费级智能眼镜无此硬件。构建纯软件替代方案需满足双重约束:预测器需足够准确以保留下游描述质量,且足够轻量以在智能手机的功耗与计算预算内运行。本文提出 EgoGazeLite,一种用于自我中心视频的轻量双进程注视预测器。在两种 MLLM、三种自动指标及两名 LLM 评判者的评估中,预测注视裁剪区域与真实注视裁剪区域无显著差异,10 种场景均证实等价性。EgoGazeLite 仅含 1570 万参数、6.71 GFLOPs,在消费级加速器硬件上,注视与裁剪全流程端到端实时运行(21.6 毫秒/帧)。这些结果使基于 MLLM 的 token 高效、注视条件自我中心视频理解无需眼动追踪硬件。

英文摘要

The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget. Memory and compute cost scale with the number of visual tokens, and high-resolution video quickly becomes expensive to transmit and process at scale. Prior work (GazeLLM) addresses this by cropping the video around the camera wearer's gaze. This reduces the number of visual tokens by about tenfold while maintaining or improving the quality of full-resolution descriptions. However, this compression strategy depends on dedicated eye-tracking hardware, which is unavailable on consumer smart glasses. Building a software-only substitute poses a joint constraint: the predictor must be accurate enough to preserve downstream description quality, yet light enough to run on-device, within the power and compute budget of a smartphone. We address this with EgoGazeLite, a lightweight dual-process gaze predictor for egocentric video. Across two MLLMs, three automated metrics, and two LLM judges, predicted-gaze crops show no significant difference from ground-truth-gaze crops. Equivalence is confirmed in all ten cases. EgoGazeLite achieves this at 15.7M parameters, 6.71 GFLOPs, and runs the full gaze-and-crop pipeline end-to-end in real time (21.6 ms/frame) on consumer accelerator hardware. Together, these results remove the need for eye-tracking hardware for token-efficient, gaze-conditioned egocentric video understanding with MLLMs.

Comments16 pages. Accepted at the WearableAI Workshop, ECCV 2026 (Archival Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑