COSMI:多物体交互的组合式合成
COSMI: COmpositional Synthesis of Multi-object Interactions
浏览论文内容
中文总结 AI 辅助
COSMI通过组合单物体交互片段,生成多物体交互数据集,并训练扩散Transformer模型,在未见物体和交互组合上实现泛化,优于基线。
中文摘要 AI 辅助
人类-物体交互的生成模型受限于现有数据:日常活动涉及多个物体,但大多数捕获数据集一次只记录一个物体,因为多物体捕获在组合上代价高昂。我们的观察是,交互是局部的,因此单物体捕获已经包含了多物体活动的组成部分。我们组合它们:接触一致的单交互片段,镜像以平衡双手,在身体间转移,以及一个语言模型和几何检查只接受在语义和物理上合理的配对。因此,数据集随片段组合增长,而非记录时间。COSMI数据集包含222k个序列和275小时,最多五个物体,几乎是最大的多物体捕获的三十倍,并且可以通过添加数据集甚至手-物体记录来扩展。在此数据上,我们训练了COSMI方法,一个文本到交互的扩散Transformer,其遵循数据的构建方式:权重共享的物体槽生成可变数量的物体,相对于移动它们的身体部位进行预测。在一个包含未见物体和未见交互组合的基准上,基于该数据集训练的模型泛化到未见组合。COSMI在文本对齐和接触准确性上优于基线,其中在未见物体上的优势最大。代码、模型和数据集流程将在项目页面上发布:此https URL。
英文摘要
Generative models of human-object interaction are bounded by the data that exists: everyday activities involve several objects, but most captured datasets record one at a time, as multi-object capture is combinatorially expensive. Our observation is that interactions are local, so single-object captures already contain the parts of multi-object activities. We compose them: contact-consistent clips of single interactions, mirrored to balance the hands, transfer between bodies, and a language model and geometric checks admit only the pairings that are plausible, semantically and physically. Therefore, the dataset grows combinatorially with the clips rather than recording time. The COSMI dataset holds 222k sequences and 275 hours with up to five objects, nearly thirty times the largest multi-object capture, and can be extended by adding datasets or even hand-object recordings. On this data we train the COSMI method, a text-to-interaction diffusion transformer that follows how the data is built: weight-shared object slots generate a variable number of objects, predicted relative to the body parts that move them. On a benchmark with an unseen object and unseen interaction combinations, models trained on the dataset generalize to the unseen combinations. COSMI outperforms baselines in text alignment and contact accuracy, where its margin is largest on the unseen object. Code, models, and the dataset pipeline will be released on the project page: https://ptrvilya.github.io/cosmi.
发表机构
- University of Tübingen(图宾根大学)
- Tübingen AI Center(图宾根人工智能中心)
- Zuse School ELIZA(Zuse ELIZA 学院)
- Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。