arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PortraitAes:意图条件化的结构化人像美学评估

PortraitAes: Intent-Conditioned Structured Portrait Aesthetics Assessment

Junzhou Xie, Haozhong Xiong, Xunyun Tian, Kaile Du, Tianchen Yu, Qiang Li, Wei Liu, Jiaming Liu, Ruihua Huang, Yang Shi, Guangcan Liu

arXiv 2610.05010首次发表:更新:

发表机构

Southeast University; Qwen Business Unit of Alibaba(东南大学; 阿里巴巴通义千问业务部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PortraitAes,一种基于摄影意图的结构化人像美学评估方法,通过11K基准和专家评分标准分解任务,在多任务训练及分数校准后,在标准与困难基准上显著超越现有通用和专门模型。

AI 中文摘要

人像美学评估根据人像图像实现其摄影意图的有效程度来分配可比较的分数。这些分数支持图像生成流程中的数据筛选、候选选择和偏好建模。现有方法通常预测单一美学分数,或使用不基于摄影意图进行条件化的通用多模态大语言模型(MLLMs)。这种遗漏很重要,因为相同的模糊、姿势、光照或构图选择可能服务于一种摄影意图,却损害另一种。因此,这些模型学习的是与上下文无关的美学先验,并对具有不同摄影目标的人像产生不一致、不准确且具有误导性的判断。我们引入了PortraitAes-Bench,一个11K规模的基准,将该任务分解为意图条件化的子判断。专家撰写的评分标准定义了九种摄影意图、六个一级维度和22个二级标准。它们支持一个用于意图路由、专家评估、验证和分数融合的结构化流程。遵循这一结构,我们使用多任务监督训练了PortraitAes。随后,我们通过高斯分数校准和维度内跨图像排序来提高分数的可比性。在标准基准上,PortraitAes实现了0.924的皮尔逊相关系数和0.934的斯皮尔曼秩相关系数。在困难案例集上,其皮尔逊相关系数为0.829,斯皮尔曼秩相关系数为0.795。在这两个集合上,PortraitAes均优于所评估的通用MLLMs和专门的美学基线。

英文摘要

Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Existing methods typically predict a single aesthetic score or use general-purpose MLLMs without conditioning on photographic intent. This omission matters because the same blur, pose, lighting, or framing choice may serve one photographic intent but undermine another. These models thus learn context-agnostic aesthetic priors and yield inconsistent, inaccurate, misleading judgments for portraits with distinct photographic objectives. We introduce PortraitAes-Bench, an 11K-scale benchmark that decomposes this task into intent-conditioned subjudgments. Expert-authored rubrics define nine photographic intents, six first-level dimensions, and 22 secondary criteria. They support a structured pipeline for intent routing, specialist assessment, verification, and score fusion. Following this structure, we train PortraitAes with multi-task supervision. We then improve score comparability through Gaussian score calibration and within-dimension cross-image ranking. On the standard benchmark, PortraitAes achieves a Pearson correlation of 0.924 and a Spearman rank correlation of 0.934. On the hard-case set, its Pearson correlation is 0.829 and its Spearman rank correlation is 0.795. Across both sets, PortraitAes outperforms the evaluated general-purpose MLLMs and specialized aesthetic baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑