arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01479cs.CV

CameraEditor:基于视频先验序列建模的相机控制图像编辑

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

发表机构西安交通大学 · 新加坡科技研究局
查看机构详情
  • Xi’an Jiaotong University(西安交通大学)
  • Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo

首次发表
浏览论文内容

中文总结 AI 辅助

CameraEditor是将相机控制编辑转化为时间序列预测任务的图像编辑框架,通过几何感知与动态参考机制实现高精度相机控制,在自建数据集和CamEditor-Bench上表现优于现有方法。

中文摘要 AI 辅助

除语义内容外,相机参数对决定任意图像的几何视角和外观起着关键作用。尽管近期的图像编辑模型在语义和风格操作方面表现出色,但它们在显式相机参数控制方面存在不足。在处理大视角偏移时,指令驱动模型面临两难困境:要么出现结构撕裂,要么生成忽略几何指令的保守输出。为解决该问题,我们引入CameraEditor,这一框架将相机控制编辑从空间问题重新表述为时间序列预测任务。通过利用视频扩散模型的时间一致性,我们的方法将显式几何感知模块与动态参考路由机制相结合,这使我们能够通过动态全景裁剪构建几何上严格的视觉参考对,从而克服基于文本指令的模糊性。此外,CameraEditor策略性地插入中间过渡帧以分解大视角偏移,提供强大的时间缓冲,以保留内容身份和空间一致性。我们构建了包含5760个实例的训练数据集。作为独立贡献,我们推出CamEditor-Bench,这是一个与模型无关的评估套件,包含462个测试用例。大量实验表明,CameraEditor实现了最先进的相机控制精度和源身份保留,优于现有方法。

英文摘要

Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at semantic and stylistic manipulation, they struggle with explicit camera parameter control. When handling large perspective shifts, instruction-driven models face a dilemma: they either suffer from structural tearing or generate conservative outputs that ignore geometric instructions. To address this, we introduce CameraEditor, a framework that reformulates camera-controlled editing from a spatial problem into a temporal sequence prediction task. By leveraging the temporal coherence of video diffusion models, our approach integrates an explicit geometric perception module with a dynamic reference routing mechanism. This allows us to construct geometrically rigorous visual reference pairs via dynamic panorama cropping, overcoming the ambiguity of text-based instructions. Furthermore, CameraEditor strategically inserts intermediate transition frames to decompose large perspective shifts, providing a robust temporal buffer that preserves content identity and spatial coherence. We construct a training dataset of 5,760 instances. As an independent contribution, we introduce CamEditor-Bench, a model-agnostic evaluation suite of 462 test cases. Extensive experiments demonstrate that CameraEditor achieves state-of-the-art camera control precision and source identity preservation, outperforming existing methods.

补充信息

↑