arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GenStream:面向以人为中心媒体生成式重建的语义流式传输框架

GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media

Emanuele Artioli, Daniele Lorenzi, Shivi Vats, Farzad Tashtarian, Christian Timmerer

arXiv 2609.18634首次发表:更新:

发表机构

Alpen-Adria-Universitaet(阿尔卑斯-亚得里亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GenStream提出用骨骼关键点、相机参数和静态3D背景等元数据替代视频帧传输,经生成模型重建,实现99.9%以上带宽压缩,为以人为中心媒体流式传输开辟新方向。

AI 中文摘要

视频流媒体占据了全球互联网流量的大部分,然而传统流水线对于结构化、以人为中心的内容(如体育、表演或交互式媒体)仍然效率低下。标准编解码器重新编码整个帧,包括前景和背景,对所有像素进行统一处理,忽略了场景的语义结构。这导致显著的带宽浪费,尤其是在背景静态且运动仅限于少数显著演员的场景中。我们提出了GenStream,一种语义流式传输框架,用紧凑的结构化元数据替代密集的视频帧。GenStream不是传输像素,而是将每个场景编码为骨骼关键点、相机视点参数和静态3D背景模型的组合。这些元素被传输到客户端,在客户端,生成模型重建逼真的人形,并将其从原始视点合成到3D场景中。这种范式实现了极端压缩,与HEVC相比,连续数据流的带宽减少了99.9%以上。我们在奥运会花样滑冰片段上部分验证了GenStream,并展示了在最小数据下实现高感知保真度的潜力。虽然承认转移到客户端的显著计算成本和泛化方面的挑战,GenStream为体积化身合成、跨视图的规范3D演员融合以及个性化观看体验开辟了新方向,为后编解码器时代的可扩展智能流媒体奠定了基础。

英文摘要

Video streaming dominates global internet traffic, yet conventional pipelines remain inefficient for structured, human-centric content such as sports, performance, or interactive media. Standard codecs re-encode entire frames, foreground and background alike, treating all pixels uniformly and ignoring the semantic structure of the scene. This leads to significant bandwidth waste, particularly in scenarios where backgrounds are static and motion is constrained to a few salient actors. We introduce GenStream, a semantic streaming framework that replaces dense video frames with compact, structured metadata. Instead of transmitting pixels, GenStream encodes each scene as a combination of skeletal keypoints, camera viewpoint parameters, and a static 3D background model. These elements are transmitted to the client, where a generative model reconstructs photorealistic human figures and composites them into the 3D scene from the original viewpoint. This paradigm enables extreme compression, achieving over 99.9% bandwidth reduction compared to HEVC for the continuous data stream. We partially validate GenStream on Olympic figure skating footage and demonstrate potential for high perceptual fidelity under minimal data. While acknowledging the significant computational costs shifted to the client and challenges in generalization, GenStream opens new directions in volumetric avatar synthesis, canonical 3D actor fusion across views, and personalized viewing experiences, laying the groundwork for scalable, intelligent streaming in the post-codec era.

Comments9 pages. Published at ACM MM 2025. Code: https://github.com/emanuele-artioli/genstream

Journal refIn Proceedings of the 33rd ACM International Conference on Multimedia 2025 (MM '25). ACM, New York, NY, USA, 12276-12284

DOI:10.1145/3746027.3758153

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑