Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs
用你所说的语言推理:为VideoLLM中的推理蒸馏进行个人习语式的离线策略轨迹改写
机构 * KAIST(韩国科学技术院) ; Korea University(高丽大学) ; Google Cloud AI Research(谷歌云人工智能研究院)
专题命中 其他视频模型 :video reasoning(abstract);分类 cs.CV
AI总结 该研究针对VideoLLM的推理蒸馏提出Echo-GRPO框架,通过改写教师特权轨迹为学生自身习语实现策略对齐,实例化的VideoEcho-R1在多骨干和基准测试中均获性能提升,且作为插件模块适配多种训练框架。
Comments Work in progress