arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14811cs.CV

成本落于何处:面向循环一致对抗网络稳定性增强的部署感知采用顺序

Where the Cost Falls: A Deployment-Aware Adoption Order for Stability Enhancements to Cycle-Consistent Adversarial Networks

Rowan Hussein, Mohamed Ouf

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对循环一致对抗网络的四种稳定性增强方法,提出了部署感知的采用顺序,分析了各增强的成本差异与适配场景,通过马到斑马转换实验验证了效果并明确了后续研究方向。

中文摘要 AI 辅助

采用循环一致对抗网络进行非配对图像到图像转换的团队会遇到相同的障碍:对抗训练振荡或崩溃,循环一致性保留粗略布局而精细纹理漂移,单个判别器判断全局真实性会遗漏局部伪影。四种增强方法可解决这些故障,且通常仅在输出质量上进行比较。本文表明,它们在成本落于何处方面也存在显著差异,这种差异源于架构而非特定运行,为计算或延迟预算有限的团队提供了采用顺序。带梯度惩罚的Wasserstein目标、循环重建上的VGG19感知损失以及多尺度判别器仅改变训练过程,因此团队可采用或放弃它们而不改变部署内容。自注意力单独保留在部署的生成器中,内存随特征图大小的平方增长,这使其成为资源受限团队应推迟采用的组件。我们将所有四种增强整合到轻量调优的基线模型中用于马到斑马的转换,在固定控制下逐个引入后再组合,对每种增强给出其针对的故障模式及整合方式。我们记录了基线模型产生的崩溃和重建伪影模式,报告了对保存样本进行视觉检查得到的各变体结果,以及组合模型的Fréchet Inception距离和Kernel Inception距离。我们明确了仍需的协议,涵盖各变体、感知相似性和下游分割,以基于测量证据对这些增强进行排名。

英文摘要

Teams that adopt cycle-consistent adversarial networks for unpaired image-to-image translation meet the same obstacles: adversarial training oscillates or collapses, cycle consistency preserves coarse layout while finer texture drifts, and a single discriminator judging global realism misses local artifacts. Four enhancements address these failures, and they are usually compared on output quality alone. We show that they also divide sharply by where their cost falls, and that this division, which follows from the architecture and not from any particular run, yields an adoption order for teams under a compute or latency budget. A Wasserstein objective with gradient penalty, a VGG19 perceptual loss on the cycle reconstruction, and multi-scale discriminators change training only, so a team can adopt or drop them without altering what ships. Self-attention alone persists into the deployed generator, with memory growing as the square of the feature-map size, which makes it the one component a resource-constrained team should defer. We integrate all four onto a lightly tuned baseline for horse-to-zebra translation, introduced one at a time on a fixed control and then combined, and for each we give the failure mode it targets and how it integrates. We document the collapse and reconstruction-artifact modes the baseline produced, report what visual inspection of saved samples showed for each variant, and report Fréchet Inception Distance and Kernel Inception Distance for the combined model. We specify the protocol still needed, covering the individual variants, perceptual similarity, and downstream segmentation, to rank these enhancements on measured evidence.

发表机构

  • University of Ottawa(渥太华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑