arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24107cs.CVcs.AI

MatReplace:用于室内场景材料替换的无参考、条件对齐基准

MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

发表机构新加坡国立大学 · 南洋理工大学 · VinUniversity
查看机构详情
  • National University of Singapore(新加坡国立大学)
  • Nanyang Technological University(南洋理工大学)
  • VinUniversity

机构由 AI 辅助整理,请以论文原文为准。

Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy, See-Kiong Ng, Luu Anh Tuan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对室内场景材料替换任务推出无参考基准MatReplace,定义三个条件轨道,实验发现命名材料渲染基本解决但像素视觉定位仍存挑战,专家评分验证其有效性。

中文摘要 AI 辅助

材料替换是一种常见的室内设计操作:在保持选定表面的几何形状、周围环境和光照的同时,更改其材料。尽管该任务具有商业重要性,但目前没有公开基准专门针对该任务,且对其进行评估颇具挑战性。基于参考的指标会惩罚此固有一对多设置中的有效输出,偏向参考生成器的风格,且无法公平比较接收不同形式指导的编辑器。我们推出MatReplace,这是一种无参考基准,可沿四个可验证维度评估编辑结果:局部材料正确性、全局光照协调性、外部环境保留和内部结构保留。它定义了三个轨道,每次仅改变一个条件信号:(A)仅指令、(B)指令加区域掩码、(C)用材料参考图像替代指令。我们的结果揭示了材料命名与视觉 grounding( grounding 译为“视觉定位”)之间存在明显差距。在轨道A中,领先的闭源编辑器实现了示例级别的材料渲染,且在我们的主要聚合指标下超过了示例锚点。在轨道B中,掩码仅对与掩码兼容、场景保留能力较弱的模型有帮助,任务配对的单种子效应在对齐的模型家族中从+0.137到-0.090不等。在轨道C中,参考图像条件会使所有模型家族在两种聚合指标下均出现性能下降,降幅为-0.031至-0.508;在最坏情况下,模型会重新绘制参考图像本身,且性能比返回输入不变时更差。因此,在该分布下,最强闭源编辑器已基本解决了命名材料的渲染问题,但从像素中对材料进行视觉定位仍是一个未解决的挑战。专家评分验证了我们的排名(Kendall's tau = 0.68),且与我们的聚合指标的一致性高于基于GT参考或CLIP的基线。

英文摘要

Material replacement is a common interior-design operation: changing the material of a selected surface while preserving its geometry, surroundings, and illumination. Despite its commercial relevance, no public benchmark isolates this task, and evaluating it is challenging. Reference-based metrics penalize valid outputs in this inherently one-to-many setting, favor the style of the reference generator, and cannot fairly compare editors that receive different forms of guidance. We introduce MatReplace, a reference-free benchmark that evaluates edits along four verifiable dimensions: local material correctness, global lighting harmony, outside preservation, and inside structure. It defines three tracks that vary one conditioning signal at a time: (A) instruction only, (B) instruction plus region mask, and (C) material reference image instead of instruction. Our results reveal a clear divide between naming and visually grounding materials. In Track A, leading closed-source editors achieve exemplar-level material rendering and surpass the exemplar anchor under our primary aggregate. In Track B, masks help only mask-compatible models with weak scene preservation, with task-paired, single-seed effects ranging from +0.137 to -0.090 across aligned model families. In Track C, reference-image conditioning degrades every family under both aggregates, by -0.031 to -0.508; in the worst cases, models repaint the reference image itself and perform worse than returning the input unchanged. Thus, named-material rendering is largely solved by the strongest closed editors on this distribution, but grounding materials from pixels remains an open challenge. Expert ratings validate our ranking (Kendall's tau = 0.68) and align with our aggregates more closely than GT-referenced or CLIP-based baselines.

↑