VISTA:一种基于注意力的多智能体强化学习架构,用于空间态势感知传感器任务分配
VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking
另 1 家 · 查看机构详情
- ETSIAE-School of Aeronautics, Universidad Politécnica de Madrid(马德里理工大学航空学院)
- Indra Sistemas S.A.(英德拉系统公司)
- Escuela de Ingeniería de Fuenlabrada, Universidad Rey Juan Carlos(胡安卡洛斯国王大学丰拉夫拉达工程学院)
- Department of Computer Systems Engineering, Universidad Politécnica de Madrid(马德里理工大学计算机系统工程系)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
VISTA提出一种基于注意力与指针解码的可扩展多智能体强化学习架构,使观测和动作空间独立于目录规模,显著提升空间态势感知传感器任务分配效率与泛化能力。
中文摘要 AI 辅助
在轨物体数量的快速增长正日益增加空间态势感知传感器任务分配的复杂性,对经典优化方法提出了挑战,因为这些方法需要在日益庞大的目录中分配有限、异构且分布式的感知资源。现有的深度强化学习方法在简化场景中显示出潜力,但固定维度的状态和动作表示限制了它们扩展到大型动态目录和分布式感知网络的能力。我们提出了VISTA(可变实体智能传感器任务分配架构),这是一种可扩展的深度强化学习架构,用于在可变物体群体和传感器配置下进行持续的、由不确定性驱动的目录维护。VISTA将基于物理和任务信息的top-K检索与实体中心注意力、循环记忆和基于指针的动作解码相结合,从而使每个智能体的观测和动作空间独立于目录大小。我们在不同场景下评估了VISTA,从固定大小的单传感器基准测试到大规模天基任务分配和异构协作感知。在30个轨道目标的情况下,VISTA比固定维度的循环基线快31.2%恢复目录。在大规模场景中,相对于最强的经典参考方法,VISTA将五小时不确定性降低了97.5%,相对于循环学习器降低了99.3%。对多达20,000个物体的零样本测试揭示了感知能力、目录大小和恢复时间范围之间的近线性关系。学习到的策略还表现出传感器模态适应以及对物体群体和初始不确定性变化的泛化能力。这些结果共同表明,VISTA为在大型分布式异构地基和天基传感器网络中实现自适应空间态势感知传感器任务分配提供了一个可扩展的框架。
英文摘要
The rapid growth of resident space objects is increasing the complexity of space situational awareness sensor tasking, challenging classical optimization methods as they allocate finite, heterogeneous, and distributed sensing resources across ever-larger catalogues. Existing deep reinforcement learning approaches show promise in reduced settings, but fixed-dimensional state and action representations limit their ability to scale to large, dynamic catalogues and distributed sensing networks. We introduce VISTA (Variable-Entity Intelligent Sensor Tasking Architecture), a scalable deep reinforcement learning architecture for persistent uncertainty-driven catalogue maintenance across variable object populations and sensor configurations. VISTA combines physics- and mission-informed top-K retrieval with entity-centric attention, recurrent memory, and pointer-based action decoding, thereby keeping each agent's observation and action spaces independent of catalogue size. We evaluate VISTA across different scenarios, from fixed-size single-sensor benchmarks to large-scale space-based tasking and heterogeneous cooperative sensing. With 30 orbiting targets, VISTA recovers the catalogue 31.2% faster than the fixed-dimensional recurrent baseline. In the large-scale regime, VISTA reduces five-hour uncertainty by 97.5% relative to the strongest classical reference and by 99.3% relative to the recurrent learner. Zero-shot tests up to 20,000 objects reveal near-linear relations between sensing capacity, catalogue size, and recovery horizon. Learned policies also exhibit sensor modality adaptation and generalization to population and initial-uncertainty shifts. Together, these results demonstrate that VISTA provides a scalable framework for adaptive space situational awareness sensor tasking across large, distributed networks of heterogeneous ground- and space-based sensors.