arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成野生:面向野生动物的个体一致图像到视频生成

Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife

Yuzhuo Li, Di Zhao, Xinyu Zhang, Daniel Wilson, Yun Sing Koh

arXiv 2610.05587首次发表:更新:

发表机构

University of Auckland(奥克兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对野生动物个体识别数据稀缺问题,提出高频引导的I2V框架WildIcon,通过频率感知身份编码和轻量级适配保持个体一致性,并构建数据集WildlifeVid,实验验证其优于现有方法。

AI 中文摘要

个体层面的野生动物识别常常面临数据稀缺的问题,因为同一动物在不同姿态、视角和运动下的观测数据很少。图像到视频(I2V)生成通过从单张参考图像合成额外的观测数据,为缓解这一限制提供了一种有前景的方法。然而,现有的I2V模型主要强调全局布局、语义和运动,因此往往无法保留区分不同野生动物个体的细粒度局部外观线索,如皮毛纹理、条纹边界、斑点构型和轮廓过渡。我们观察到这些身份关键线索与高频信息密切相关。为应对这一挑战,我们提出了WildIcon,一种高频引导的I2V框架,用于野生动物个体一致性。具体而言,WildIcon引入了一个频率感知的身份编码分支,从参考图像中提取个体特定的高频线索。结合隔离的前景信息,所得的身份令牌随后被注入交叉注意力块作为身份条件。基于冻结的主干网络和轻量级身份适配,WildIcon在保留参考图像中可见的细粒度身份线索的同时,保持了基础I2V模型的运动可控性和语义保真度。此外,为支持野生动物个体一致I2V的训练和评估,我们构建了WildlifeVid,一个以野生动物为中心的视频数据集,包含高质量、时间连贯的片段和个体级别的身份标签。在I2V生成和下游动物重识别(ReID)上的实验表明,WildIcon相比现有基线实现了更强的个体一致性,且其过滤后的输出可作为下游ReID的有用候选训练增强数据。

英文摘要

Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a single reference image. However, existing I2V models mainly emphasize global layout, semantics, and motion, and therefore often fail to preserve fine-grained local appearance cues that distinguish one wildlife individual from another, such as fur texture, stripe boundaries, spot configurations, and contour transitions. We observe that these identity-critical cues are closely related to high-frequency information. To address this challenge, we propose WildIcon, a high-frequency-guided I2V framework for wildlife individual consistency. Specifically, WildIcon introduces a frequency-aware identity encoding branch that extracts individual-specific high-frequency cues from the reference image. Combined with isolated foreground information, the resulting identity tokens are then injected into cross-attention blocks as identity conditioning. Building on a frozen backbone with lightweight identity adaptation, WildIcon preserves fine-grained identity cues visible in the reference image while retaining the motion controllability and semantic fidelity of the base I2V model. In addition, to support the training and evaluation of wildlife individual-consistent I2V, we construct WildlifeVid, a wildlife-centric video dataset with high-quality, temporally coherent clips and individual-level identity labels. Experiments on I2V generation and downstream animal re-identification (ReID) show that WildIcon achieves stronger individual consistency than existing baselines, and that its filtered outputs can serve as useful candidate training augmentations for downstream ReID.

Comments29 pages, 12 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑