AI 中文总结
UniQuery4R是一种查询条件4D场景重建框架,通过单编码多帧片段实现灵活的源-目标选择,在WorldTrack数据集的场景流估计和运动点重建任务中获最佳宏平均结果。
AI 中文摘要
重建动态4D场景需要联合估计对应关系、几何结构、物体运动和相机运动。现有前馈方法通常预测密集的任务特定映射,或独立处理源-目标对,这会给稀疏查询带来不必要的计算,且不同帧对之间的特征复用受限。我们提出UniQuery4R,一种查询条件框架,该框架对多帧片段仅编码一次,仅在解码阶段通过源-目标交叉注意力选择源视图、目标视图和连续源图像坐标。每个查询联合预测目标对应关系、目标时刻3D位置、场景流以及源深度,相机参数按每个视图估计。该设计允许编码后的片段在任意源-目标选择中复用,通过批量查询支持稀疏推理和密集重建,无需与固定片段长度绑定的学习型时间嵌入。我们进一步引入场景流的方向-幅值参数化,对运动点和静态点分别进行监督。在所有评估方法中,UniQuery4R在WorldTrack数据集上的场景流估计和运动点重建任务中均取得了最佳的宏平均结果。
英文摘要
Reconstructing dynamic 4D scenes requires jointly estimating correspondence, geometry, object motion, and camera motion. Existing feed-forward methods typically predict dense task-specific maps or independently process source-target pairs, leading to unnecessary computation for sparse queries and limited feature reuse across different frame pairs. We present UniQuery4R, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention. Each query jointly predicts target correspondence, target-time 3D position, and scene flow, along with source depth, while camera parameters are estimated per view. This design allows the encoded clip to be reused across arbitrary source-target selections and supports both sparse inference and dense reconstruction through batched queries, without learned temporal embeddings tied to a fixed clip length. We further introduce a direction-magnitude parameterization of scene flow with separate supervision for moving and static points. Among the evaluated methods, UniQuery4R achieves the best macro-average results on WorldTrack for both scene-flow estimation and dynamic-point reconstruction.
Comments16 pages, 8 figures. Project page: https://kosmoresearch.github.io/UniQuery4R/