arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越单一视频:电商跨视频推理的基准构建与主动证据寻求

Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng

arXiv 2610.03099首次发表:更新:

发表机构

Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对电商跨视频推理难题,提出首个基准AdsCVR和智能体框架AdSeek,通过主动证据获取与离线轨迹修正提升性能,在测试集上大幅超越基线。

AI 中文摘要

电商视频信息密集,消费者在评估产品、商家在评估营销策略时经常对视频进行比较。然而,现有的多模态模型主要聚焦于单一视频理解,在跨视频信息比较方面能力有限。我们提出了AdsCVR,这是首个电商跨视频推理基准,包含2,483个视频和6,110个问答对,覆盖六个推理维度。跨视频推理要求模型在大量冗余帧中定位细粒度证据,并整合视觉细节、语音和屏幕文本。为此,我们提出了AdSeek,一个智能体框架,在多轮探索中动态选择视觉和音频工具,用主动证据获取取代静态均匀采样。为了解决强化学习中稀疏的信用分配问题,我们开发了一种离线轨迹修正机制,用于识别强化学习生成轨迹中的推理错误和缺失的多模态证据。修正后的轨迹提供监督微调信号,减少强化学习期间学习到的偏差。该机制支持一个修正的自举流程,其中初始强化学习暴露推理瓶颈,监督微调纠正这些瓶颈,最终强化学习阶段进一步改进策略。AdSeek在AdsCVR测试集上达到74.30%的准确率,比其Qwen3-VL-8B-Instruct基础模型高出27.90个百分点。它还能泛化到开放领域的CrossVid基准,展示了有效的主动证据收集能力。

英文摘要

E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. Cross- video reasoning requires models to locate fine-grained evidence among many redundant frames and integrate visual details, speech, and on-screen text. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. To address the sparse credit assignment of reinforcement learning, we develop an offline trajectory rectification mechanism that identifies reasoning errors and missing multimodal evidence in RL-generated trajectories. The corrected trajectories provide supervised fine-tuning signals that reduce biases learned during RL. This mechanism supports a rectified bootstrapping pipeline in which initial RL exposes reasoning bottlenecks, supervised fine-tuning corrects them, and a final RL stage further improves the policy. AdSeek achieves 74.30 percent accuracy on the AdsCVR test split, outperforming its Qwen3-VL-8B-Instruct backbone by 27.90 percentage points. It also generalizes to the open- domain CrossVid benchmark, demonstrating effective active evidence gathering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑