AI 中文总结
研究开放词汇对象目标导航中智能体状态更新问题,提出基于线性注意力的导航策略LANav,采用线性注意力作主干,引入加权状态扩展线性注意力提升性能,在多方面优于基于Transformer基线,证明其可行性与跨场景鲁棒性。
AI 中文摘要
开放词汇对象目标导航(OVON)要求智能体在部分可观测性下运行,有效的内部状态更新对导航性能至关重要。此更新由策略网络实现,近期方法采用基于Transformer的主干,通过上下文窗口的自注意力整合时间信息。然而,控制实验表明基于Transformer的策略下性能不随上下文长度扩展,质疑自注意力在导航中状态整合的适用性。为此,提出基于线性注意力的导航(LANav),采用线性注意力作为策略主干以保持结构化状态更新。在相同设置下评估的多个LANav变体始终优于基于Transformer的基线。随着状态更新机制更结构化和规范,性能提升,凸显状态更新设计的重要性。为提高状态更新有效性,引入加权状态扩展线性注意力(WSLA),将每个注意力头的状态扩展为多个子状态并使用可学习加权读出聚合扩展子状态。配备WSLA的LANav在HM3D - OVON上实现36.4%的平均成功率,在宏观平均成功率上比基于Transformer的对应方法高6.3个百分点,同时保持计算效率。距离分层结果显示在长距离情节中有更大提升,HSSD转移和微调证明跨场景分布的鲁棒性。在Unitree Go2上的实际部署在50次试验中进一步实现82%的成功率,支持LANav的实际可行性和从模拟到现实的转移。
英文摘要
Open-Vocabulary Object Goal Navigation (OVON) requires agents to operate under partial observability, making effective internal state updates critical for navigation performance. This update is implemented by the policy network, where recent approaches adopt Transformer-based backbones with self-attention over a context window to integrate temporal information. However, our controlled experiments show that performance does not scale with context length under Transformer-based policies, questioning the suitability of self-attention for state integration in navigation. To this end, we propose Linear Attention-based Navigation (LANav), which adopts linear attention (LA) as the policy backbone to maintain a structured state update rather than self-attention over the context window. Across multiple LA variants evaluated under identical settings, LANav consistently outperforms Transformer-based baselines. Performance improves as state update mechanisms become more structured and regulated, highlighting the importance of state update design. To improve state update effectiveness, we introduce Weighted State-Expansion Linear Attention (WSLA), which expands each attention head's state into multiple sub-states and uses learnable weighted readout to aggregate expanded sub-states. Equipped with WSLA, LANav achieves 36.4% average success rate (SR) on HM3D-OVON, outperforming Transformer-based counterparts by 6.3 percentage points in macro-averaged SR, while maintaining computational efficiency. Distance-stratified results show larger gains in long-distance episodes, while HSSD transfer and fine-tuning demonstrate robustness across scene distributions. Real-world deployment on a Unitree Go2 further achieves an 82% success rate over 50 trials, supporting the practical feasibility and sim-to-real transfer of LANav.
Comments12 pages, 7 figures