发表机构
Alibaba Group; Beijing University of Posts and Telecommunications(阿里巴巴集团; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频美学评估问题,提出受峰值-结尾规则启发的Peak-End-Net框架,通过引入预训练IAA头部、设计美学节奏编码器和动态门控融合机制,基于冻结ViT实现,在实验中取得最优性能。
AI 中文摘要
视频美学评估(VAA)旨在预测视频的美学吸引力,但与其他视觉评估任务相比,其探索较少。其进展受到大规模基准稀缺以及美学判断内在主观性的阻碍。本文从心理学角度重新审视VAA,提出了受峰值-结尾规则启发的轻量级且可解释的框架Peak-End-Net。通过引入预训练的图像美学评估(IAA)头部来生成逐帧美学先验,设计美学节奏编码器以及动态门控融合机制,该方法基于冻结的视觉Transformer(ViT),参数少且可扩展。在两个现有VAA基准上的大量实验表明其达到了当前最优性能。
英文摘要
Video aesthetic assessment (VAA) aims to predict how aesthetically pleasing a video is, yet remains far less explored than other visual assessment tasks. Its progress is hindered not only by the scarcity of large-scale benchmarks, but also by the intrinsic subjectivity of aesthetic judgment, which is shaped by human perception. In this paper, we revisit VAA from a psychological perspective and propose \textit{Peak-End-Net}, a lightweight and interpretable framework inspired by the \textit{peak-end rule}, which suggests that people tend to judge a temporal experience mainly according to its salient moments and the ending. Building on this intuition, we first transfer knowledge from image aesthetic assessment (IAA) to VAA by introducing a pretrained IAA head to produce frame-wise aesthetic priors, which serve as surrogate signals for identifying aesthetically salient moments and guiding \textit{peak-end rule}-based temporal aggregation. To further capture how a video evolves aesthetically over time, we design an aesthetic rhythm encoder that models temporal progression beyond isolated moments. Additionally, we refine the overall assessment through a dynamic gated fusion mechanism to improve robustness under distribution shift. Our method is built on a frozen vision transformer (ViT) and requires only a small number of trainable parameters, making it scalable and parameter-efficient. Extensive experiments on two existing VAA benchmarks, including in-domain evaluation on VADB and cross-domain testing on DIVIDE-3K, demonstrate that our approach achieves state-of-the-art performance, affirming the value of psychologically grounded modeling for VAA. Our code and models are available at https://github.com/AMAP-ML/Peak-End-Net.
CommentsAccepted to ACM MM 2026, Code: https://github.com/AMAP-ML/Peak-End-Net