AI 中文总结
针对城市社会语义分割中目标边界划分不准确的问题,提出EVEREST模型,通过以自我为中心的探索策略、伪代码规范执行逻辑及强化学习,在真实数据集上实现最优性能。
AI 中文摘要
城市社会语义分割利用数字和卫星图像,为城市资源分配等下游应用提供关键的空间语义信息。尽管现有方法已实现较高的分割精度,但仍存在目标边界划分不准确的问题,其根本原因在于当前模型主要依赖被动聚合的全局跨模态线索,缺乏对环境的主动探索。为解决这一局限,我们提出EVEREST模型,采用以自我为中心的探索策略,使模型能够主动探查边界线索并进行自我修正;此外,我们将离散自然语言提示表述为伪代码,以规范执行逻辑,进一步采用强化学习实现这一不可约过程并激发模型的结构化推理能力。我们的EVEREST在真实世界城市社会语义数据集的所有指标上均取得最优性能,证明了该模型的优越性,代码可在该https链接获取。
英文摘要
Urban socio-semantic segmentation leverages digital and satellite imagery to provide critical spatial semantic information for downstream applications such as urban resource allocation. Although existing methods achieve high segmentation accuracy, they still suffer from inaccurate delineation of target boundaries. The underlying issue is that current models primarily rely on passively aggregated global cross-modal cues, lacking active exploration of the environment. To address this limitation, we propose the EVEREST model, which adopts an egocentric exploration strategy that enables the model to actively investigate boundary cues and perform self-correction. In addition, we formulate discrete natural-language prompts as pseudocode to regularize the execution logic. Reinforcement learning is further employed to implement this irreducible process and elicit the model's structured reasoning capability. Our EVEREST achieves optimal performance on all metrics in the real world urban socio-semantic dataset, demonstrating the superiority of our model. Codes are available at https://github.com/TechCloud-x/EVEREST.