用于多摄像机视角推荐的双Transformer模型
A Dual-Transformer for Multi-Camera View Recommendation
浏览论文内容
中文总结 AI 辅助
提出带交叉注意力的双Transformer模型,在TVMCE数据集上的Precision@0.5达56.60%,优于SOTA,用20%视频微调可实现高效的剪辑风格个性化。
中文摘要 AI 辅助
多摄像机系统是现代媒体制作的基础,多摄像机剪辑是一项关键任务,涉及在每个时刻正确选择合适的摄像机视角。本文提出了一种新颖的带交叉注意力(Cross-Attention)的双Transformer架构,在TVMCE数据集(电视节目多摄像机剪辑数据集)上的性能大幅优于当前的SOTA模型。我们的模型将任务解耦为两个部分:(1)专用的时间编码器首先处理过去帧的序列,以构建丰富的近期历史记忆;(2)候选摄像机视角随后作为查询,通过交叉注意力模块作用于该记忆,使每个候选视角能独立查询历史上下文,找到对自身评估最相关的信息。我们的方法达到了56.60%的Precision@0.5,相比之前的最佳结果37.16%有显著提升。我们进一步进行了消融研究,探索使用轻量级骨干网络架构,其中SwinV2骨干网络取得了最佳性能,达到69.65%的Precision@0.5。使用这一最佳配置,我们随后研究了调整模型以复制特定人类剪辑师的剪辑风格的可行性。为此,我们使用目标视频初始片段的不同比例对模型进行微调。结果表明,即使仅使用视频的20%进行微调,模型的Precision@0.5也出现了可测量的提升,显示出在针对每个电视节目或制作人进行数据高效的剪辑风格个性化方面的巨大潜力。
英文摘要
Multi-camera systems are foundational to modern media production, and multi-camera editing is a critical task. This involves the proper selection of the appropriate camera view at each moment. In this paper, we propose a novel Dual-Transformer architecture with Cross-Attention that heavily outperformed the current SOTA models over the TVMCE dataset (TV Shows Multicamera Editing dataset). Our model decouples these tasks: (1) a dedicated temporal encoder first processes the sequence of past frames to build a rich memory of the recent history, and (2) the candidate camera views then act as queries to this memory via a cross-attention module, allowing each candidate to independently interrogate the historical context and find the most relevant information for its own evaluation. Our approach achieved 56.60% Precision@0.5, representing a substantial improvement over the prior best result of 37.16%. We further conducted an ablation study exploring the use of lightweight backbone architectures, where the SwinV2 backbone yielded the best performance, achieving 69.65% Precision@0.5. Using this best-performing configuration, we then investigated the feasibility of adapting the model to replicate the editing style of a specific human editor. To this end, we fine-tuned the model using varying proportions of the initial segment of a target video. Our results demonstrate that even with only 20% of the video used for fine-tuning, the model exhibited measurable improvements in Precision@0.5, indicating strong potential for data-efficient personalization of editing style adapted to each individual TV show or producer.
发表机构
- Universitat de Barcelona(巴塞罗那大学)
机构由 AI 辅助整理,请以论文原文为准。