发表机构
INRIA; LIPADE, Université Paris Cité; DM3L, University of Zurich; CIRAD(法国国家信息与自动化研究所; 巴黎西岱大学LIPADE实验室; 苏黎世大学DM3L实验室; 法国农业发展研究国际合作中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Atomizer-IO架构,以原子表示和局部交叉注意力取代网格假设,在多种传感任务上媲美专用模型,并自然扩展至无序点云。
AI 中文摘要
大多数视觉架构假设观测数据位于规则网格上,这种抽象对自然图像有效,但对于通道、时间采样、空间分辨率和几何形状可能变化的传感数据而言则具有限制性。通用的基于集合的架构去除了网格,但也去除了有用的空间归纳偏置。我们引入了Atomizer-IO,这是一种将观测置于首位并从其物理关系中推导结构的架构。基于数据的原子表示,每个观测由其测量值和采集元数据描述,而局部交叉注意力将观测映射到可任意放置的锚点。我们通过逐步放宽网格假设来评估这一设计,从变化的输入栅格配置和不完整的通道集合,到灵活的输出密度,最终到没有栅格网格的输入。Atomizer-IO在大多数任务上与灵活的EO专用架构具有竞争力,同时提供训练后对推理成本的控制以及有竞争力的计算-性能权衡。相同的公式无需架构重新设计即可扩展到无序的3D点云,表明原子接口可推广到常规栅格输入之外。这些结果表明,像素、块和网格不必定义传感架构的接口。
英文摘要
Most vision architectures assume that observations lie on a regular grid, an effective abstraction for natural images but a restrictive one for sensing data whose channels, temporal sampling, spatial resolution, and geometry can vary. Generic set-based architectures remove the grid, but also remove useful spatial inductive biases. We introduce Atomizer-IO, an architecture that places observations first and derives structure from their physical relationships. Building on top of an atomic representation of the data, each observation is described by its measurement and acquisition metadata, while local cross-attention maps observations to anchor points that can be arbitrarily placed. We evaluate this design by progressively relaxing the grid assumption, from varying input raster configurations and incomplete channel sets to flexible output density and, ultimately, inputs without a raster grid. Atomizer-IO is competitive with flexible EO-specific architectures on most tasks, while offering post-training control over inference cost and competitive compute--performance trade-offs. The same formulation extends without architectural redesign to unordered 3D point clouds, showing that the atomic interface generalizes beyond regular raster inputs. These results suggest that pixels, patches, and grids do not need to define the interface of a sensing architecture.