arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效推理蒸馏:通过合成思维链和难度感知微调实现小型视频语言模型

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania

arXiv 2609.16255首次发表:更新:

发表机构

Liverpool John Moores University; Carnegie Mellon University; Meta; University of Texas at Arlington(利物浦约翰摩尔大学; 卡内基梅隆大学; Meta; 德克萨斯大学阿灵顿分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种高效推理蒸馏方法,通过约900个不确定性筛选示例和4B教师生成的合成思维链微调2B视频语言模型,在极低计算成本下超越4倍规模模型,并发现将思维链置于答案后可显著提升紧凑模型推理能力。

AI 中文摘要

我们提出了一种高效的方法,将推理能力蒸馏到紧凑的视频语言模型(VLM)中,用于视频问答(VideoQA)。我们的方法仅使用约900个基于不确定性筛选的示例对2B参数模型进行微调,每个示例都附带了由4B教师模型生成的合成思维链(CoT)理由。尽管计算成本极低——在单个A100 GPU上不到两小时——我们的方法使2B模型能够超越高达4倍规模的VLM,并在CinePile、ActivityNet-QA和MLVU上实现泛化,接近其自身4B教师模型的性能。一个关键发现是,将CoT理由放在答案之后——与标准提示相反——显著提高了紧凑模型的推理能力。这一见解挑战了流行的CoT惯例,并揭示了在有限模型容量下的新对齐策略。我们的发现为训练适合移动和边缘应用的可部署、推理丰富的VLM提供了实用蓝图。

英文摘要

We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Our approach fine-tunes a 2B-parameter model using only $\sim$900 uncertainty-selected examples, each augmented with synthetic chain-of-thought (CoT) rationales generated by a 4B teacher. Despite its minimal compute cost - under two hours on a single A100 GPU - our method enables the 2B model to outperform VLMs up to 4$\times$ larger, and generalize across CinePile, ActivityNet-QA, and MLVU, approaching the performance of its own 4B teacher. A key finding is that placing CoT rationales after the answer - contrary to standard prompting - substantially improves reasoning in compact models. This insight challenges prevailing CoT conventions and reveals new alignment strategies under limited model capacity. Our findings offer a practical blueprint for training deployable, reasoning-rich VLMs suited for mobile and edge applications.

Comments14 pages, 2 figures, 5 tables. Published in MultiMedia Modeling (MMM 2026), LNCS 16412

Journal refMultiMedia Modeling (MMM 2026), Lecture Notes in Computer Science, vol. 16412, pp. 567-580, Springer, Singapore, 2026

DOI:10.1007/978-981-95-6950-2_40

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑