RetailSMV:零售场景中基础视频世界模型的外视角与内视角适应
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
浏览论文内容
中文总结 AI 辅助
研究零售场景下基础视频世界模型的外视角与内视角适应,发现仅用外视角数据训练的模型在多项指标上优于或等同于联合训练。
中文摘要 AI 辅助
基础视频扩散模型越来越被视为具身智能体的世界模拟器,但其在互联网规模通用视频上的预训练使其与真实世界部署领域对齐不佳。我们研究将预训练的基础视频世界模型参数高效地适应到零售场景:当同一活动的同步内视角和外视角视频可用时,哪种视角的训练数据能产生最强的适应模型?我们引入了RetailSMV(零售同步多视角),一个包含来自五个超市的32,105个带字幕零售片段的语料库,这些片段从商店员工视角(上架、整理、称重、管理供应车、结账扫描)同步捕获内/外视角,而非先前零售视频语料库以顾客为中心,并在相同超参数下训练了三个匹配的Cosmos3-Nano低秩适应(LoRA)配置(仅内视角、仅外视角、联合)。在一个200片段保留测试集上,使用七种互补指标在严格配对统计协议下评估,仅外视角适应在七项点估计中的六项上匹配或超过联合适应,并在LPIPS、PSNR和DreamSim上显著更好,尽管仅使用15,985个外视角片段(联合为32,105)训练。对称配对比较进一步表明,将外视角数据添加到仅内视角训练中有帮助,而将内视角数据添加到仅外视角训练中有害。绝对适应差距在最短展开时间处最大,将近水平预测窗口确定为适应最有益的区间。
英文摘要
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the strongest adapted model? We introduce RetailSMV (Retail Synchronized Multi-View), a corpus of 32,105 captioned retail clips from five supermarkets with synchronized ego/exo capture from the store-staff perspective (stocking, arranging, weighing, managing supply carts, scanning at checkout), rather than the customer-centric framing of prior retail video corpora, and train three matched Low-Rank Adaptation (LoRA) configurations of Cosmos3-Nano (egocentric-only, exocentric-only, combined) under identical hyperparameters. On a 200-clip held-out test set evaluated with seven complementary metrics under a strict paired statistical protocol, exocentric-only adaptation matches or exceeds combined adaptation on six of seven point estimates and is significantly better on LPIPS, PSNR, and DreamSim, despite training on only 15,985 exocentric clips (versus 32,105 for combined). A symmetric paired comparison further shows that adding exocentric data to egocentric-only training helps while adding egocentric data to exocentric-only training hurts. The absolute adaptation gap is largest at the shortest rollout time, identifying the near-horizon prediction window as the regime in which adaptation is most beneficial.
发表机构
- DreamVu
机构由 AI 辅助整理,请以论文原文为准。