arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MPIE-Bench:解剖学上合理的多人交互编辑基准测试

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie

arXiv 2607.27616首次发表:更新:

发表机构

University of Science and Technology of China; Metastone Technology(中国科学技术大学; 元石科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多人交互编辑的解剖学与几何错误问题,构建了MPIE-Bench基准,提出MPIE-Eval评估方法,发现现有编辑模型在两个维度表现不均衡,且其评估比VLM清单更贴合人类判断。

AI 中文摘要

文本到图像和个性化编辑模型如今能轻松合成高保真的单主体图像,但将多个指定人物放入拥抱、搬运、搏斗等共享接触动作时,仍会出现肢体融合、生成多余四肢、身体穿透等严重错误。现有评估大多忽略这些解剖学和几何学问题,而视觉语言模型(VLM)作为评判者的清单常聚焦于“是否存在交互”,错误却对人类显而易见。我们推出MPIE-Bench,这是一个包含2500个样本的基准,由视频挖掘的编辑三元组构成,涵盖405个场景、14类交互以及4种接触密度(C0-C3)。我们还提出MPIE-Eval,其两个新维度从冻结的公开多人网格重建中对接触时的几何结构评分:解剖学维度检查是否有完整的重建人体来解释每个人类状质量;交互维度检查人体间的穿透和表面距离是否符合指令要求的接触。在10种编辑模型中,两种不同模型的网格解剖学评分最高为0.65,网格交互评分最高为0.72,因此没有任何一种编辑模型在两个维度上都表现出色,而VLM清单对相同图像的评分超过0.95。一项五人评判研究证实,这两个维度比零样本VLM评判更贴合人类判断,且在去除所有权重和阈值后排名依然稳定。

英文摘要

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑