发表机构
Virginia Tech; Amazon(弗吉尼亚理工大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对产品视频属性值提取,提出无需训练的ViS-CoT流水线,结合视觉聚类、搜索和交错思维链推理,在VideoAVE数据集上使多个视频VLM平均提升17.91个百分点micro-F1。
AI 中文摘要
现有的视觉属性值提取(AVE)方法主要依赖于静态产品图像,无法捕捉时间线索、多角度视图和细粒度的视觉细节。直接将视频视觉语言模型(VLM)应用于产品AVE会因缺乏领域知识而导致性能有限,而对它们进行微调则需要大量高质量数据和可观的计算资源。因此,我们提出了视觉搜索增强的思维链推理(ViS-CoT),这是一种无需训练、即插即用的流水线,可以轻松应用于任何开源视频VLM,用于电子商务中的视频到文本AVE。具体来说,ViS-CoT采用视觉聚类来识别代表性帧,随后通过视觉搜索检索语义相似的产品知识,以丰富属性线索。接下来,一个交错的CoT推理模块通过由字幕生成和自动语音识别得到的视觉对齐辅助文本,迭代地细化推理。最后,整合的信息引导模型做出准确且细粒度的属性预测。在VideoAVE数据集上跨越14个产品类别的广泛实验表明,ViS-CoT持续增强了多个最先进的视频VLM,在micro-F1上实现了平均17.91个百分点的提升。
英文摘要
Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failing to capture temporal cues, multi-angle views and fine-grained visual details. Directly applying video vision-language models (VLMs) to product AVE results in limited performance due to the lack of domain knowledge, and fine-tuning them requires extensive high-quality data and substantial computational resources. Thus, we propose visual search augmented chain-of-thought reasoning (ViS-CoT), a training-free, plug-and-play pipeline that can be easily applied to any open-source video VLM for video-to-text AVE in e-Commerce. Specifically, ViS-CoT employs visual clustering to identify representative frames, followed by visual search to retrieve semantically similar product knowledge that can enrich attribute cues. Next, an interleaved CoT reasoning module iteratively refines reasoning through visually-aligned auxiliary texts derived from captioning and automatic speech recognition. Finally, the integrated information guides the model toward accurate and fine-grained attribute predictions. Extensive experiments across 14 product categories on the VideoAVE dataset show that ViS-CoT consistently enhances multiple state-of-the-art video VLMs, achieving an average improvement of 17.91 percentage points in micro-F1.
Comments17 pages, 6 figures, accepted for publication in EMNLP 2026 Findings