具身多媒体:教程
Embodied Multimedia: A Tutorial
浏览论文内容
中文总结 AI 辅助
本教程提出具身多媒体跨学科研究范式,构建四层统一架构,明确五大前沿应用方向,探讨开放挑战,为具身智能背景下的多媒体领域研究提供指引。
中文摘要 AI 辅助
传统多媒体技术围绕为人类观察者优化内容传输构建,从感知驱动的压缩标准到以人类为中心的质量指标。随着具身智能的快速发展,自主智能体必须在物理世界中实时感知、推理和行动,这暴露出传统多媒体基础设施与具身任务需求之间的根本不匹配。就此,本教程论文正式将具身多媒体作为一门跨学科研究范式提出,该范式将多模态数据视为贯穿感知-决策-行动全循环的感知与通信基础。具体而言,我们提出了包含数据层、通信层、认知层和评估层的四层统一架构,并对各层内的关键支撑技术进行了结构化综述。此外,我们明确了具身多媒体有望发挥基础性作用的五个前沿应用方向:多媒体通信、物理智能、具身异常感知、元宇宙与交互式多媒体,以及AI驱动的艺术创作。本文还探讨了开放技术挑战与未来研究方向,以指导该新兴领域的研究群体。
英文摘要
Traditional multimedia technology has been built around optimizing content delivery for human observers, from perceptually driven compression standards to human-centric quality metrics. With the rapid rise of embodied intelligence, autonomous agents must perceive, reason, and act within the physical world in real time, exposing fundamental mismatches between conventional multimedia infrastructure and the demands of embodied tasks. In this regard, this tutorial paper formally introduces Embodied Multimedia as a cross-disciplinary research paradigm that treats multimodal data as the perceptual and communicative substrate spanning the full perception-decision-action loop. To be specific, we present a four-layer unified architecture comprising Data, Communication, Cognitive, and Evaluation layers, and provide a structured review of key enabling technologies within each layer. Furthermore, we identify five frontier application directions where Embodied Multimedia is positioned to serve a foundational role: multimedia communication, physical intelligence, embodied anomaly perception, the metaverse and interactive multimedia, and AI-driven art creation. Open technical challenges and future research directions are discussed to guide the community in this emerging field.
发表机构
- Tongji University(同济大学)
- Fudan University(复旦大学)
- Simon Fraser University(西蒙菲莎大学)
- University of Ottawa(渥太华大学)
机构由 AI 辅助整理,请以论文原文为准。