发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该综述首次以统一框架系统研究智能眼镜,提出L0-L5框架等,连接任务与多类要素,制定评估协议,为智能眼镜的可信第一人称智能发展提供路线图。
AI 中文摘要
智能眼镜正从捕获与显示配件演变为连接人类感知、持续上下文及数字或物理行动的第一人称智能平台。其佩戴视角与佩戴者的视觉、听觉、运动及手物交互相契合,但必须在严格的能量、热、隐私和反馈约束下运行。尽管增强现实、自我中心视觉、多模态模型、人机交互和具身智能领域进展迅速,但现有文献在设备、任务和基准方面仍呈碎片化状态。关键挑战不在于模型能否单独进行识别、回答、记忆或行动,而在于完整系统能否维持可靠、时间有效、可修正且可管控的感知-状态-交互-行动循环。本综述是首个通过此类统一框架系统研究智能眼镜的工作。我们将第一人称数据流和约束任务效用形式化,沿八个可验证硬件能力轴对设备进行表征,围绕七个相互依存的基础能力组织文献,并引入覆盖捕获、反应感知、上下文辅助、持续状态、受控行动和具身耦合的L0-L5框架。在九个应用场景中,我们将任务与数据集、系统、产品、利益相关者、失败后果及证据缺口相连接。我们还提出了九维部署框架、声明条件评估协议,以及从受控测量到纵向实地验证和审计的证据阶梯。这些元素共同使智能眼镜更具可比性、可部署性和可重复评估性,同时勾勒出通往可信第一人称智能的路线图。
英文摘要
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.
CommentsProject at https://github.com/zhangzjn/awesome-smart-glasses