arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26620cs.CV

GeoComposer:基于几何的摄影构图指令

GeoComposer: Geometry-Grounded Photographic Composition Instruction

Shuangzhi Li, Qiaoqiao Jia, Xingxin Chen, Guile Wu, Dongfeng Bai

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出GeoComposer框架,通过几何感知表示学习和混合奖励强化学习,实现基于三维几何一致性的摄影构图编辑,生成视觉美观且几何一致的构图范例。

中文摘要 AI 辅助

摄影构图旨在为改善图像的取景、视角和空间排列提供视觉指导。早期方法主要依赖图像裁剪来增强构图,这受限于输入图像的视角和空间排列。近期方法探索了图像理解与编辑以改善构图,但主要关注指令遵循和美学质量,忽视了三维场景几何一致性对摄影构图的重要性。在本工作中,我们提出GeoComposer,一种新颖的基于几何的摄影构图框架,该框架分析给定图像的构图以生成文本指导,并合成一个增强给定图像构图的视觉范例。为促进基于几何的构图,我们提出一种几何感知表示学习机制,利用视觉几何基础模型中的几何先验来塑造构图编辑模型的中间表示。该机制保留了全局结构关系和局部细粒度对应关系,以实现基于几何的构图。此外,我们提出一种由混合奖励引导的强化学习策略,该策略联合优化指令遵循、美学质量和几何一致性。这使得模型能够生成忠实遵循构图指令同时保持视觉吸引力和几何一致性的视觉范例。大量实验表明,我们的方法优于现有最先进方法,突显了其在生成视觉吸引且几何一致的构图方面的有效性。

英文摘要

Photographic composition aims to provide visual guidance for improving the framing, viewpoint, and spatial arrangement of an image. Early methods primarily rely on image cropping to enhance composition, which is restricted to the viewpoint and spatial arrangement of the input image. Recent methods have explored image understanding and editing to improve composition, but they mainly focus on instruction following and aesthetic quality, overlooking the importance of 3D scene geometry consistency for photographic composition. In this work, we propose GeoComposer, a novel geometry-grounded photographic composition framework that analyzes the composition of a given image to generate textual guidance and synthesizes a visual exemplar that enhances the composition of the given image. To promote geometry-grounded composition, we propose a geometry-aware representation learning mechanism that leverages geometric priors from a visual geometry foundation model to shape the intermediate representations of the composition editing model. This mechanism preserves both global structural relationships and local fine-grained correspondences for geometry-grounded composition. Furthermore, we propose a reinforcement learning strategy guided by a hybrid reward that jointly optimizes instruction following, aesthetic quality, and geometric consistency. This enables the model to generate visual exemplars that faithfully follow the composition instructions while remaining visually appealing and geometrically consistent. Extensive experiments show the superiority of our approach over state-of-the-art methods, highlighting its effectiveness in generating visually appealing and geometrically consistent composition.

发表机构

  • University of Alberta(阿尔伯塔大学)
  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

↑