发表机构
East China Normal University; Peking University; Sun Yat-sen University; Tencent(华东师范大学; 北京大学; 中山大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AESOP提出统一的不对称框架,通过独立人体路径和共享相机生成器,并利用轨迹对学习平移强度条件,在PulpMotion数据集上实现高质量相机生成与强度控制。
AI 中文摘要
人体运动定义了动作,而相机轨迹决定了动作的呈现方式。针对给定人体运动的相机生成以及联合的人体-相机生成通常被视为独立任务,尽管两者共享一种不对称依赖关系:人体运动可以独立生成,而相机则响应于已实现的动作。我们提出了AESOP,一个统一框架,包含独立的人体路径和共享的以人体为条件的相机生成器。其不对称架构同时服务于两个任务,同时在相机生成过程中保持人体输出不变。尽管人体上下文将镜头锚定到动作,且相机文本描述了其运动,但平移强度仍未得到充分指定。因此,我们构建了在相机平移幅度上不同但共享相同人体运动和相机文本的轨迹对,然后利用这些对来学习显式的强度条件。在PulpMotion数据集上的实验表明,在两个任务中均展现出强大的相机分布和构图质量,以及对相机平移强度的有效控制。
英文摘要
Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.