AI 中文总结
本文针对公开图像编辑API无法提供模型内部信息的问题,提出客户端通过掩码合成可实现整容预览的区域定位,在保证身份保留的同时提升编辑区域的准确性,且效果优于仅用提示词的方式。
AI 中文摘要
若向商业图像编辑器请求整容手术预览,其往往会对脸部进行超出请求范围的修改:例如修改鼻子时可能同时平滑皮肤或调整光线。现有的将编辑限制在单一区域的方法需要访问模型内部信息,而公开的编辑API并不提供此类信息。本文探究仅从客户端可实现的控制程度。在一项试点基准测试中,6种商业编辑配置和1种基于掩码的修复模型分别执行面部拉皮式颌颈修改及隆鼻术修改,设置三种客户端控制级别:仅通过提示词;通过地标生成的掩码将编辑区域从响应中裁剪并粘贴回原始照片(掩码合成);在模型支持的情况下要求模型在掩码内进行修复。共尝试210次编辑,其中196次可评分。采用ArcFace余弦相似度衡量身份保留程度;采用CIELAB像素变化率衡量变化落在请求区域而非受保护面部区域的比例。在12张正面人脸的区域指标可评分的情况下,掩码合成相较于配对的仅提示词输出,定位效果的中位数提升0.446(95%人脸聚类自举区间为0.421-0.457),同时对请求区域的修改程度相近。不同编辑器在编辑强度与身份保留之间存在差异,且本文测试的1种修复模型未优于简单合成方法。相对于每张人脸的输入到术后基线,没有编辑器的输出在身份嵌入层面更接近术后照片。本研究聚焦于控制,而非临床准确性:无外科医生对输出进行评分,且每种条件仅生成一次。在该研究范围内,将整容预览限制在预期区域内无需访问模型;客户端的掩码与合成操作可在所有测试的编辑器上实现该效果,且对提供商成本较低。
英文摘要
Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smooth skin or alter lighting. Existing methods for confining an edit to one region require access to the model's internals, which a public editing API does not expose. We ask how much control is possible from the client side alone. In a pilot benchmark, six commercial editing configurations and one mask-based inpainting model perform facelift-style jaw-neck and rhinoplasty edits at three levels of client-side control: the prompt alone; cutting the edited region out of the response and pasting it back onto the original photograph through a landmark-derived mask (a masked composite); and asking the model itself to inpaint inside the mask where supported. Of 210 attempted edits, 196 could be scored. ArcFace cosine measures identity preservation; a CIELAB pixel-change ratio measures how much change lands inside the requested region rather than a protected facial zone. On the 12 frontal faces the regional metric could score, the masked composite improved localization over the paired prompt-only output by a median of 0.446 (95% face-clustered bootstrap interval 0.421-0.457) while changing the requested region about as much. Editors differed in edit strength versus identity retention, and the one inpainting model we tested did not beat the simple composite. Against each face's input-to-postoperative baseline, no editor moved its outputs closer to the postoperative photograph in identity-embedding terms. This is a study of control, not clinical accuracy: no surgeons rated the outputs, and each condition was generated once. Within that scope, keeping a surgical preview inside its intended region needs no access to the model; a mask and composite on the client enforce it across every editor tested, at low provider cost.
Comments9 pages, 7 figures, 2 tables. Pilot study; no surgeon ratings. Code and paper source: https://github.com/suxrobGM/localize-dont-beautify