发表机构
Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对历史街景生成中记录缺失问题,提出CrossTimeEdit模型,利用近期街景与时间卫星差异进行编辑式生成,结合SFT与RL优化,在三个编辑标准上较基线提升17.12%。
AI 中文摘要
历史街景图像记录了城市演变,但不均匀的覆盖范围在历史记录中留下了大量空白。生成可信的过去外观需要恢复已改变的结构,同时保留持久的场景内容。我们构建了VIGOR-his,一个跨十年的跨视角数据集,包含三大洲11个城市的43,653个位置级四元组。其自动化流程包括空间配对、一致性筛选、变化分类,以及基于卫星的变化描述和局部编辑指令的生成与验证。基于VIGOR-his,我们提出了CrossTimeEdit,一种将历史街景生成重新表述为编辑的模型,利用近期街景来约束视角和不变外观,并将时间卫星差异作为变化证据。从FLUX.2 [Klein] 4B出发,我们通过监督微调(SFT)和在线强化学习(RL)训练CrossTimeEdit。我们设计了三个街景编辑标准,即指令对齐(IA)、背景保留(BP)和质量与物理合理性(QP),既作为RL奖励维度,也作为评估指标。我们使用流匹配模型的组内相对策略优化(Flow-GRPO)结合组奖励解耦归一化策略优化(GDPO)来优化这一多奖励目标,该策略在聚合前对每个奖励维度进行归一化。CrossTimeEdit在三个编辑标准上的整体性能相比预训练基线提升了17.12%,并在场景一致性、视觉真实感和感知质量上优于跨视角生成模型。实现代码、数据集和模型权重可在以下网址获取:https URL。
英文摘要
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at https://luhanwen67.github.io/CrossTimeEdit-release/.
Comments43 pages, 9 figures, 7 tables. Project page: https://luhanwen67.github.io/CrossTimeEdit-release/