几何驱动的合成人脸阴影协调:一种乘性、保持反照率的重光照管线
Geometry-Driven Shadow Harmonisation for Composited Faces: A Multiplicative, Albedo-Preserving Relighting Pipeline
AI总结:
针对换脸结果光度不一致问题,提出几何驱动的乘性阴影注入重光照管线,仅通过变暗保持反照率与色调,避免重绘人脸,实验验证了其有效性与参数影响。
AI中文摘要:
换脸和人脸合成管线通常会产生几何对齐良好但光度上不合理的人脸:供体人脸带有平坦、近正面的影棚光照,而宿主身体和背景则带有方向性场景光。大多数现有修复方法通过颜色迁移、神经重光照或逆渲染重新合成人脸,因此存在改变身份、肤色和纹理的风险。我们提出一种保守的替代方案:几何驱动的形状阴影注入。该管线从不重新绘制人脸。它从光栅化的3D人脸代理估计逐像素增益场$g\in[g_{\min},1]$,并将其按通道均匀地乘到线性RGB上,因此该算子只能变暗而不能改变色度。密集地标网格被光栅化为深度缓冲区,从中我们导出表面法线、空腔项和屏幕空间投射阴影。主光方向从宿主侧线索(身体、背景、头发光晕)估计;人脸侧线索被降权,因为它们会恢复供体的光照。阴影幅度不与宿主匹配:它由三参数传递$(\tau,\sigma,g_{\min})$设置。着色场除以皮肤上的第75百分位数,然后在羽化的、皮肤门控的人脸掩膜内进行门控、缩放、钳位、平滑和重新裁剪。在解析人脸高度场上,默认$(\tau,\sigma,g_{\min})=(0.90,0.45,0.82)$修改了56.5%的人脸像素,平均增益为0.938(修改像素上为0.890),并将3.4%的像素驱动到下限。色调不变性是算子的推论。我们以闭式形式分析该传递,消融其参数,并讨论单调、仅变暗公式的失败模式,包括非平坦供体的双重阴影。
英文摘要:
Face swapping and face compositing pipelines routinely produce a face that is geometrically well aligned but photometrically implausible: the donor face carries flat, near-frontal studio illumination while the host body and background carry directional scene light. Most existing remedies re-synthesise the face through colour transfer, neural relighting, or inverse rendering, and therefore risk altering identity, skin tone, and texture. We present a conservative alternative: geometry-driven form-shadow injection. The pipeline never repaints the face. It estimates a per-pixel gain field $g\in[g_{\min},1]$ from a rasterised 3D face proxy and multiplies it channel-uniformly onto linear RGB, so the operator can only darken and cannot shift chromaticity. A dense landmark mesh is rasterised into a depth buffer, from which we derive surface normals, a cavity term, and screen-space cast shadows. Key-light direction is estimated from host-side cues (body, background, hair halo); on-face cues are downweighted because they recover the donor's lighting. Shadow magnitude is not matched to the host: it is set by a three-parameter transfer $(τ,σ,g_{\min})$. The shading field is divided by its 75th percentile over skin, then gated, scaled, clamped, smoothed, and re-clipped inside a feathered, skin-gated face mask. On an analytic face heightfield, the default $(τ,σ,g_{\min})=(0.90,0.45,0.82)$ modifies 56.5% of face pixels with mean gain 0.938 (0.890 on modified pixels) and drives 3.4% of pixels to the floor. Hue invariance is a corollary of the operator. We analyse the transfer in closed form, ablate its parameters, and discuss failure modes of a monotone, darkening-only formulation, including double-shadowing of non-flat donors.