arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Crayotter:通过组相对偏好反向传播学习长时程视频编辑智能体

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation

Lecheng Yan, Jianze Lin, Yichong Zhang, Ben Pan, Wenxi Li, Chenyang Lyu, Liting Zhou, Cathal Gurrin

arXiv 2608.02694首次发表:更新:

发表机构

University of Science and Technology of China; Beijing Normal University; Jilin University; Tianjin University; East China Normal University; Alibaba Group; Dublin City University(中国科学技术大学; 北京师范大学; 吉林大学; 天津大学; 华东师范大学; 阿里巴巴集团; 都柏林城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 GRPB 方法,构建长时程视频编辑任务集,训练出的 9B Crayotter 模型在 AgenticVBench 上超越多个专有系统,实现从主观延迟结果中学习视频编辑智能体。

AI 中文摘要

长时程视频编辑智能体仅在做出多个相互依赖的决策后才能获得最终产品反馈。然而,编辑质量具有主观性,存在多种有效解决方案,且在不同请求间无法进行有意义的校准,这使得全局标量目标既模糊又缺乏时间维度的信息。我们的关键发现是,固定请求、素材和制作约束可将该主观目标转化为直接可比替代方案间的序数比较。我们提出组相对偏好反向传播(GRPB),其将同一任务的排名转化为零和优势,并将其作为有界信用重新分配到语义编辑片段中。滞后分配器和防护传输机制可防止当前判断或不可靠估计直接影响同一 rollout 组。我们手动构建了一套项目不相交、时长分层的真实编辑任务,用于训练和受控评估。在匹配基线、信用干预、外部基准测试和盲法人类评估中,GRPB 均改善了编辑行为和渲染产品。由此得到的 9B Crayotter 模型在 AgenticVBench 上超越了多个专有系统,支持任务局部偏好简化作为从主观、延迟结果中学习的实用方法。代码及所有支持材料可在该 https URL 公开获取。

英文摘要

Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. Our key observation is that fixing the request, materials, and production constraints converts this subjective objective into an ordinal comparison among directly comparable alternatives. We introduce Group-Relative Preference Backpropagation (GRPB), which transforms same-task rankings into zero-sum advantages and redistributes them as bounded credit over semantic editing segments. A lagged allocator and guarded transmission prevent current judgments or unreliable estimates from directly shaping the same rollout group. We manually construct a project-disjoint, horizon-stratified suite of realistic editing tasks for training and controlled evaluation. Across matched baselines, credit interventions, external benchmarking, and blinded human evaluation, GRPB improves both editing behavior and rendered products. The resulting 9B Crayotter model surpasses several proprietary systems on AgenticVBench, supporting task-local preference reduction as a practical approach to learning from subjective, delayed outcomes. Code and all supporting materials are publicly available at https://github.com/idwts/Crayotter.

Comments12 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑