发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多目标跟踪问题,核心方法是用轻量级LoRA微调文本到视频扩散模型生成带持久身份颜色的视频,无需传统跟踪组件。主要贡献是在DanceTrack测试服务器上达到40.3 HOTA,有独特错误分布,颜色分配可在遮挡后重新识别身份。
AI 中文摘要
多目标跟踪(MOT)传统上分解为检测后关联,对象身份作为外部状态维护,如跟踪缓冲区、运动模型和外观嵌入。本文探讨视频生成器能否在像素中维护该状态。通过用轻量级上下文LoRA微调22B文本到视频扩散模型(LTX - 2.3),将RGB视频转换为ID映射视频,每个人被赋予持续不变的独特颜色。长视频按链式窗口生成,每个窗口以上一个窗口的清理尾部为条件。简短的延续微调使模型扩展给定颜色,之后身份无需跟踪器、运动模型和重新识别模块即可在链中流动。在DanceTrack测试服务器上,该系统达到40.3 HOTA,虽低于当前专业水平,但具有独特的错误分布:关联分数(AssA 44.1)超过原始基准套件中的所有跟踪器,检测是唯一不足。控制比较表明机制很重要:经典事后关联得分将相同生成窗口链接起来效果差2倍(18.2 HOTA),帧到帧IoU关联会分割跟踪,而生成器的颜色能保持跟踪完整。在383个挖掘的遮挡事件中,生成器在间隙后以42%的条件率重新获取身份,而外观嵌入基线得分为零,包括长于其时间上下文的间隙,证明生成器的颜色分配起到了紧急重新识别信号的作用。最后作者发布了代码、检查点和完整的预注册实验日志。
英文摘要
Multi-object tracking (MOT) is conventionally decomposed into detection followed by association, with object identity maintained as external state: track buffers, motion models, and appearance embeddings. We ask whether a video generator can maintain that state in pixels. We fine-tune a 22B text-to-video diffusion model (LTX-2.3) with a lightweight in-context LoRA to translate an RGB clip into an ID-map clip, a video in which every person is painted a flat, distinct color that persists over time: same color, same identity. Long videos are generated as chained windows, where each window is conditioned on the cleaned tail of the previous one. A brief continuation fine-tune teaches the model to extend a given coloring, after which identity flows through the chain with no tracker, no motion model, and no re-identification module. On the DanceTrack test server, our system, to our knowledge the first generative tracker evaluated there and the only entry with no detector and no tracking stack, reaches 40.3 HOTA. This is well below today's specialist state of the art (>=70 HOTA), but with a unique, inverted error profile: its association score (AssA 44.1) exceeds every tracker of the original benchmark suite while detection remains the sole deficit. Controlled comparisons show the mechanism matters: the same generated windows linked by classical post-hoc association score 2x worse (18.2 HOTA), and frame-to-frame IoU association fragments tracks that the generator's colors keep whole. On 383 mined occlusion events, the generator re-acquires identities after gaps at a 42% conditional rate where appearance-embedding baselines score zero, including gaps longer than its temporal context, evidence that the generator's color assignment functions as an emergent re-identification signal. We release code, checkpoints, and the full pre-registered experimental log.