发表机构
Huazhong University of Science and Technology; The Hong Kong University of Science and Technology (Guangzhou)(华中科技大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出ACM MM 2026单图像多视图合成挑战赛的获奖方案GoM,通过场景不相交验证等技术,在40个训练场景下实现293支队伍中排名第一,揭示小数据生成建模中验证设计与训练轨迹控制的关键作用。
AI 中文摘要
我们提出了ACM 2026年多媒体大会单图像引导多角度图像合成挑战赛的获奖解决方案。该方案在293支注册队伍中排名第一;在公开的A阶段排行榜上,有56支队伍获得了至少一次评分提交。该挑战仅提供40个训练场景,要求从一张RGB图像生成26个目标视图,且每个视图仅需一次前向传播;挑战禁止使用显式几何、外部渲染、链式生成、候选选择和后处理。我们发现了一个关键的模型选择缺陷:训练场景与验证场景重叠会使记忆效应表现为可迁移的视图控制。因此我们引入了GoM,即Generalization over Memorization(超越记忆的泛化),该框架结合了场景不相交的验证、曝光匹配的选择以及针对性的扩散适配。其合成模型对40亿参数的rectified-flow DiT进行适配,采用秩为32的LoRA、优化器重启、后期检查点平均以及VAE解码器调优。超过300次离线实验和24次在线提交表明,在小数据生成建模中,验证设计和训练轨迹控制的重要性可与架构规模相媲美。
英文摘要
We present the winning solution to the ACM Multimedia 2026 Grand Challenge on Single-Image Guided Multi-Angle Image Synthesis. It ranks first among 293 registered teams; 56 teams obtained at least one scored submission on the public Phase-A leaderboard. With only 40 training scenes, the challenge requires 26 target views from one RGB model and one forward pass per view; it prohibits explicit geometry, external rendering, chained generation, candidate selection, and post-processing. We identify a critical model-selection failure: shared training and validation scenes make memorization appear as transferable view control. We therefore introduce GoM. Short for Generalization over Memorization, the framework combines scene-disjoint validation, exposure-matched selection, and targeted diffusion adaptation. Its synthesis model adapts a 4B rectified-flow DiT using rank-32 LoRA, optimizer restarts, late-checkpoint averaging, and VAE decoder tuning. More than 300 offline experiments and 24 online submissions show that validation design and training-trajectory control can matter as much as architecture scale in small-data generative modeling.
CommentsACM Multimedia 2026 Grand Challenge Track winning paper; presents the 1st-place solution among 293 registered teams. 7 pages, 2 figures