arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Human-JEPA:一种以人类为中心的视觉模型,可感知并进行预测

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

Hui Wei, Licai Sun, Guoying Zhao

arXiv 2608.21160首次发表:更新:

发表机构

ELLIS Institute Finland; University of Oulu(芬兰ELLIS研究所; 奥卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Human-JEPA模型,以锚定预测方式在视频上训练,解决现有以人类为中心视觉模型无法实现运动与预测的问题,在姿态和人物重识别任务上参数更少且性能更优,可同时满足理解人类的两大需求。

AI 中文摘要

理解人类的机器应能感知当下并预测未来。现有的以人类为中心的视觉模型均在人类图像上进行预训练,在静态密集感知方面达到了顶尖水平,但无法实现运动与预测功能。本文提出Human-JEPA,这是一种以人类为中心的视觉模型,通过锚定预测的方式在视频上进行训练:将密集目标固定在初始化的冻结副本上,防止密集感知的静默崩溃;同时将块掩码替换为纯过去到未来的划分,避免了5点动作分类和17点重识别的崩溃。在冻结探针设置下,Human-JEPA在姿态和人物重识别任务上,以仅2.7倍更少的参数数量领先于基于像素锚定的专用模型,仅在高分辨率密集解析方面略有让步;其发布的预测器头部是首个不会降低预测性能的模型。因此,一个经安全适配的单一模型即可满足理解人类的两大需求。

英文摘要

Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑