发表机构
Kuaishou Technology; Institute of AI for Industries, Chinese Academy of Sciences(快手科技; 中国科学院人工智能产业研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对电商直播中产品理解缺乏产品身份与时间证据关联评估的问题,提出GPUB基准和UniPro模型,实现联合产品识别与时间定位,显著提升性能。
AI 中文摘要
电商直播已成为向在线消费者展示产品的重要渠道,其中包含多个产品,其信息分散在不同时刻。这对下游产品理解应用(如以产品为中心的直播剪辑)构成了重大挑战,此类应用需要模型识别产品及其相关片段以进行信息收集。然而,现有的通用产品理解基准通常孤立地评估产品检索和时间定位,导致产品身份与时间证据之间的关键对应关系在很大程度上未被评估。为解决这一局限,我们引入了GPUB,一个大规模基准,包含3,000个直播实例,具有质量控制的多时刻时间标注,以及超过31K时尚产品的目录。GPUB支持三个评估任务:主任务Grounded Product Understanding (GPrU)要求从直播视频和候选产品集中联合识别目标产品并定位其支持时刻;Product Retrieval和Product Moment Localization作为两个互补子任务。对现有多模态模型的评估显示,GPrU仍然极具挑战性,最佳性能基线仅达到10.13%的Pair mAP@.3。为缩小性能差距,我们进一步开发了UniPro,一个统一的产品理解模型,从共享多模态编码中导出产品对齐和时间结构化的表示,将Pair mAP@.3提升至21.53%,同时在GPrU上实现37.23%的Joint R@1@.3。
英文摘要
E-commerce livestreams have emerged as an important channel for presenting products to online consumers, often featuring multiple products with relevant information distributed across different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 real-world e-commerce livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: given a livestream video and a candidate product set, the main task Grounded Product Understanding (GPrU) requires jointly identifying the product being presented and localizing its supporting moments; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing off-the-shelf baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 24.58% while achieving 38.81% Joint R@1@.3 on GPrU.