编辑器内部的相机:用绘制的标定图案读取图像编辑器的隐式相机
The Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns
浏览论文内容
中文总结 AI 辅助
通过让图像编辑器绘制棋盘格,利用消失点几何无训练地读出其隐式相机参数,发现编辑器在俯仰和焦距上优于GeoCalib,并揭示横滚和焦距先验及LoRA控制强度不足。
中文摘要 AI 辅助
基于指令的图像编辑器能够插入物体、重绘场景风格并渲染新视角,但在它们向照片中绘制内容时,我们并不知道它们假设了何种相机模型。当要求编辑器用地板棋盘格覆盖场景时,编辑器会绘制出射影结构,经典消失点几何可从该结构中直接读出俯仰角、横滚角、焦距、偏航角,以及在渲染结果中的主点位置,整个过程无需任何训练。与GeoCalib等针对图像估计相机的标定器不同,本方法隔离出编辑器在绘制时所采用的相机。在120个具有精确真值的渲染相机上,Qwen-Image-Edit-2511绘制的瓷砖边缘与其消失点的偏差在0.26度以内,其隐式相机在俯仰角上与真实相机匹配精度达到0.8度,焦距匹配精度达到6%,除横滚角外均优于GeoCalib。当要求绘制地平线或标记消失点时,编辑器反而失败,因此这一知识是通过绘制行为而非我们尝试的显式任务揭示的。隐式相机存在两个先验:横滚角被拉向水平(斜率0.71),长焦透视被拉向约30毫米的默认值,这与模型在无场景情况下绘制的相机大致吻合。对于Qwen模型,当模糊去除五分之四的线条证据时,这些先验并未增强。在真实照片上先验更强,在NYUv2数据集上,较短的任务措辞消除了横滚角方面的差异。在24–240毫米变焦镜头拍摄的照片上,绘制的透视仅以镜头斜率的0.62倍增长,而GeoCalib和MoGe-2分别在约52毫米和42毫米处饱和。FLUX.1 Kontext和LongCat-Image-Edit则受到更强的拉动。最后,对于水平相机,相机控制LoRA仅以50–70%的强度执行姿态命令,且绘制在其输出上的棋盘格与其生成的相机一致。
英文摘要
Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which estimates the camera of an image, this isolates the camera under which the editor paints. On 120 rendered cameras with exact ground truth, Qwen-Image-Edit-2511 paints tile edges that meet their vanishing points within 0.26 degrees, and its implicit camera matches the true one to 0.8 degrees in pitch and 6% in focal length, more accurately than GeoCalib except in roll. Asked to draw the horizon or mark a vanishing point instead, the editor fails, so this knowledge is revealed by painting and not by the explicit tasks we tried. The implicit camera has two priors: roll is pulled towards level (slope 0.71), and telephoto perspective towards a default of about 30 mm, which roughly matches the camera the models paint without any scene. For Qwen, the priors do not grow when blur removes four fifths of the line evidence. They are stronger on real photographs, and on NYUv2 a shorter wording of the task removes the difference for roll. On photographs from a 24--240 mm zoom lens the painted perspective grows with only 0.62 of the lens's slope, while GeoCalib and MoGe-2 saturate at about 52 and 42 mm. FLUX.1 Kontext and LongCat-Image-Edit are pulled much harder. Finally, from a level camera a camera-control LoRA executes pose commands at only 50--70% of their strength, and a board painted into its output agrees with the camera it produced.
发表机构
- University of Hagen(哈根大学)
机构由 AI 辅助整理,请以论文原文为准。