arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12517cs.CV

一种技能并不适合所有情况:长视频问答中帧选择技能的自动发现与分类引导路由

One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering

Jian Hu, Zixu Cheng, Da Li, Wei Li, Ziquan Liu, Shaogang Gong

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AutoSkill,一种自动发现和路由帧选择技能的框架,通过分类引导为不同问题分配最优技能,在长视频问答中分别提升Qwen2.5-VL-7B和Qwen3.5-4B性能2.4%和1.2%。

中文摘要 AI 辅助

长视频问答(LVQA)需要在有限的帧预算下,在小时级视频中定位关键证据。大多数免训练方法对所有问题采用相同的帧选择策略,尽管不同问题类型所需的证据存在显著差异。我们的分析表明,帧选择策略的相对有效性在不同语义类别和基准之间有所变化,这促使我们采用自适应证据获取。在本文中,我们介绍了AutoSkill,一个源监督框架,用于自动发现和路由可执行的帧选择技能。从一个小的标记源池开始,LLM代理迭代地提出、实现、评估和细化候选技能。对于目标基准,AutoSkill仅使用未标记的问题和选项文本,来归纳共享语义分类法,将标记的源示例改写为目标风格,并估计类别到技能的映射。此过程不使用目标视频或目标答案。在推理时,每个问题被分配一个技能,该技能选择用于冻结视频MLLM单次推理的帧。在五个长视频基准分割中,AutoSkill分别将Qwen2.5-VL-7B和Qwen3.5-4B提高了2.4%和1.2%,证明了我们AutoSkill的有效性。

英文摘要

Long-Video Question Answering (LVQA) requires locating decisive evidence in hour-scale videos under a limited frame budget. Most training-free methods apply the same frame-selection strategy to all questions, despite substantial variation in the evidence required by different question types. Our analysis shows that the relative effectiveness of frame-selection strategies varies across semantic categories and benchmarks, motivating adaptive evidence acquisition. In this paper, we introduce AutoSkill, a source-supervised framework for automatically discovering and routing executable frame-selection skills. Starting from a small labelled source pool, LLM agents iteratively propose, implement, evaluate, and refine candidate skills. For a target benchmark, AutoSkill uses only unlabelled question and option text to induce a shared semantic taxonomy, rewrite labelled source examples into the target style, and estimate a category-to-skill mapping. Neither target videos nor target answers are used in this process. At inference time, each question is assigned one skill, which selects the frames used in a single inference of the frozen video MLLM. Across five long-video benchmark splits, AutoSkill improves Qwen2.5-VL-7B and Qwen3.5-4B by 2.4% and 1.2%, respectively, demonstrating the effectiveness of our AutoSkill.

发表机构

  • Queen Mary University of London(伦敦玛丽女王大学)
  • Samsung AI Research Institute(三星人工智能研究院)
  • Nanyang Technological University, Singapore(新加坡南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑