PERSIST:用于镜头边界检测的持久状态判别
PERSIST: Persistent-State Discrimination for Shot Boundary Detection
浏览论文内容
中文总结 AI 辅助
PERSIST框架将镜头边界检测重新定义为边界语义判别,结合FiLM条件正弦表示网络与多线索结构化判别器,在严格训练下大幅降低误报,且在多类数据集上达到最优性能。
中文摘要 AI 辅助
镜头边界检测(SBD)通常被视为局部视觉不连续性的定位,但手持抖动、光照闪烁、运动模糊、遮挡以及受损档案材料等因素会产生同样剧烈的局部变化,却不会引入新镜头,从而导致大量误报。我们将SBD重新表述为边界语义判别:仅当某帧的局部变化证据伴随视频潜在时间状态的持久更新,而非返回周围趋势的瞬态偏移时,才将该帧判定为边界。该持久测试通过FiLM条件正弦表示网络生成的连续潜在状态,以及结合局部变化、瞬态脉冲、回归趋势三种语义线索的结构化判别器来实现,该判别器在双速率时间主干上生成可解释的逐帧信号。由此得到的框架PERSIST使每个决策都可检查:持久准则被训练到分类器中,其逐帧效果可从门控三元组中读取,学习到的潜在状态可测量为边界判别性。在包含2727个视频的按子类型划分的诊断集上,与同等训练的线索检测器相比,它消除了33%-80%的闪光、文本叠加和档案误报;在匹配真实转换召回率时,它将TransNetV2在该诊断集上的伪事件误报大致减半,在ClipShots素材上的误报减少约四分之一,同时保持召回率。此外,在严格得多的训练条件下,它在在线、广播、短视频和历史档案迁移评估中达到与最强公开检测器相当的性能:它仅从ClipShots真实转换中学习,而基准模型利用额外语料库,其转换中85%为合成数据。代码可在此https URL获取。
英文摘要
Shot boundary detection (SBD) is widely treated as the localisation of local visual discontinuities, yet many false positives such as hand-held shake, illumination flicker, motion blur, occlusion, and damaged archival material produce equally sharp local change without introducing a new shot. We reformulate SBD as boundary semantic discrimination: a frame is favoured as a boundary only when its local change evidence is accompanied by a persistent update of the video's latent temporal state, rather than a transient excursion that returns to the surrounding trend. This persistence test is operationalised with a continuous latent state from a FiLM-conditioned sinusoidal representation network and a structured discriminator that combines three semantic cues, local change, transient impulse, and return-to-trend, into a single interpretable per-frame signal over a dual-rate temporal backbone. The resulting framework, PERSIST, turns every decision into an inspectable one: the persistence criterion is trained into the classifier, its per-frame effect stays readable from the gate triple, and its learned latent state is measurably boundary-discriminative. On a 2,727-video per-subtype diagnostic it removes 33-80% of flash, text-overlay, and archival false positives relative to an identically trained cue detector, and at matched true-transition recall it roughly halves TransNetV2's pseudo-event false positives on that diagnostic and cuts its false positives on ClipShots footage by about a quarter, while preserving recall. It does so while reaching parity with the strongest public detector across online, broadcast, short-form, and historical-archive transfer evaluations, under markedly stricter training: it learns from ClipShots real transitions only, whereas the anchor draws on additional corpora whose transitions are 85% synthetic. Code is available at https://github.com/linty5/PERSIST.