LEGO:通过视觉指令微调学习以自我为中心的动作帧生成
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
查看机构详情
- GenAI, Meta(GenAI部门,Meta公司)
- Georgia Institute of Technology(佐治亚理工学院)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出新问题以自我为中心的动作帧生成,通过视觉指令微调利用VLLM增强动作描述并提供嵌入条件,显著提升了扩散模型在Ego4D和Epic-Kitchens数据集上的图像合成效果。
中文摘要 AI 辅助
从以自我为中心的视角生成人类日常动作的教学图像是实现高效技能转移的关键步骤。在本文中,我们引入了一个新问题——以自我为中心的动作帧生成。其目标是通过以用户提示和输入的以自我为中心的图像为条件,合成一张描绘用户所处情境下动作的图像(即动作帧)。值得注意的是,现有的以自我为中心的动作数据集缺乏描述动作执行的详细标注。此外,由于领域差距,现有的基于扩散模型的图像操作模型在控制以自我为中心的图像像素空间中动作的状态转换方面表现欠佳。为此,我们提出通过视觉指令微调来学习以自我为中心的(LEGO)动作帧生成。首先,我们引入了一种提示增强方案,通过视觉指令微调从视觉大语言模型(VLLM)生成丰富的动作描述。然后,我们提出了一种新方法,利用来自VLLM的图像和文本嵌入作为附加条件,以提升扩散模型的性能。我们在两个以自我为中心的数据集——Ego4D和Epic-Kitchens上验证了我们的模型。我们的实验表明,在定量和定性评估方面,相较于先前的图像操作模型均有显著提升。我们还进行了详细的消融研究和分析,以提供关于我们方法的见解。更多关于数据集和代码的详细信息可在网站获取。
英文摘要
Generating instructional images of human daily actions from an egocentric viewpoint serves as a key step towards efficient skill transfer. In this paper, we introduce a novel problem -- egocentric action frame generation. The goal is to synthesize an image depicting an action in the user's context (i.e., action frame) by conditioning on a user prompt and an input egocentric image. Notably, existing egocentric action datasets lack the detailed annotations that describe the execution of actions. Additionally, existing diffusion-based image manipulation models are sub-optimal in controlling the state transition of an action in egocentric image pixel space because of the domain gap. To this end, we propose to Learn EGOcentric (LEGO) action frame generation via visual instruction tuning. First, we introduce a prompt enhancement scheme to generate enriched action descriptions from a visual large language model (VLLM) by visual instruction tuning. Then we propose a novel method to leverage image and text embeddings from the VLLM as additional conditioning to improve the performance of a diffusion model. We validate our model on two egocentric datasets -- Ego4D and Epic-Kitchens. Our experiments show substantial improvement over prior image manipulation models in both quantitative and qualitative evaluation. We also conduct detailed ablation studies and analysis to provide insights in our method. More details of the dataset and code are available on the website (https://bolinlai.github.io/Lego_EgoActGen/).