arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

聚焦视觉焦点增强:一种意图驱动的图像修图智能体

Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement

Chujie Qin, Zilong Zhang, Zewei Chang, Chunle Guo, Ruixing Wang, Tao Hu, Ming-Ming Cheng, Chongyi Li

arXiv 2609.01148首次发表:更新:

发表机构

Nankai University; DJI Technology Co., Ltd; NKIARI; AAIS, Nankai University(南开大学; 大疆科技有限公司; NKIARI; 南开大学先进人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MLLM驱动的智能体EyeControl,结合扩散修图执行器,通过弱用户意图引导实现图像视觉焦点增强,构建了评估数据集ControlArt-Bench,实验表明其意图对齐性更强。

AI 中文摘要

图像修图通常被定义为通过色彩调整提升整体视觉质量,但在实际应用中,它还用于通过引导观众注意力指向特定主体或区域来强调视觉焦点。实现这种面向焦点的修图本质上具有挑战性,因为它需要协调的全局和局部调整,以操纵感知显著性同时保持视觉自然性。这一复杂过程通常需要大量专业知识。在本研究中,我们提出EyeControl,一种多模态大语言模型(MLLM)驱动的智能体,配备基于扩散的修图执行器,可在弱用户意图下实现视觉焦点增强。仅需几次点击或粗略笔触,EyeControl即可将视觉注意力引导至目标区域,有效“点出”图像的视觉焦点。核心思路是在修图过程中明确关联弱用户意图、目标编辑区域和对应的色调调整操作。为实现这一点,系统首先解释意图和图像内容以推断视觉焦点,并为修图执行器生成结构化意图引导;其次,鼓励修图执行器对目标区域做出更强响应,使其注意力图与设计的伪意图图明确对齐。我们还引入操作一致性约束以改进全局和局部调整之间的协调,实现更自然连贯的修图。此外,我们提供ControlArt-Bench,一个用于视觉焦点增强的高质量评估数据集。大量评估表明,EyeControl能产生感知上有吸引力的结果,且具有更强的意图对齐性。代码可在该https URL获取。

英文摘要

Image retouching is commonly formulated as enhancing overall visual quality through color adjustment, but in practice, it also serves to emphasize visual focus by guiding viewers' attention toward a specific subject or region. Achieving such focus-oriented retouching is inherently challenging, as it requires well-coordinated global and local adjustments to manipulate perceptual saliency while maintaining visual naturalness. This intricate process typically demands substantial professional expertise. In this study, we propose EyeControl, a MLLM-driven agent with a diffusion-based retouching executor that enables visual focus enhancement under weak user intent. With only a few clicks or coarse strokes, EyeControl directs visual attention to the intended region, effectively "dotting the eye" of the image. The core idea is to explicitly link the weak user intention with the target editing region and the corresponding tonal adjustment operations during retouching. To achieve this, the system first interprets the intent and image content to infer the visual focus and generate structured intent guidance for the retouching executor. Second, the retouching executor is encouraged to respond more strongly to the target region, explicitly aligning its attention map with a designed pseudo-intent map. We also introduce an operation-consistency constraint to improve coordination between global and local adjustments, achieving more natural and coherent retouching. Additionally, we contribute ControlArt-Bench, a high-quality evaluation dataset for visual focus enhancement. Extensive evaluations demonstrate that EyeControl yields perceptually appealing results with stronger intent alignment. Code can be found in https://github.com/DragonisCV/EyeControl.

CommentsAccepted to the European Conference on Computer Vision (ECCV) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑