城市区域嵌入中移动性拓扑结构与时间语义的协同融合
Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding
另 1 家 · 查看机构详情
- Urban AI Institute, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院城市人工智能研究院)
- Graduate School of Data Science, KAIST(韩国科学技术院数据科学研究生院)
- Department of Industrial Engineering and Systems, KAIST(韩国科学技术院工业工程与系统系)
- Department of Civil and Environmental Engineering, KAIST(韩国科学技术院土木与环境工程系)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出MoSS方法,融合移动数据的序列视图与拓扑结构视图,通过协同模块捕捉高阶共现信号,仅用移动数据即在城市预测任务中超越依赖辅助模态的基线。
中文摘要 AI 辅助
城市区域嵌入已在多种城市感知任务中展现出良好效果,如犯罪预测、收入预测和服务呼叫预测。近期方法通过将移动数据与辅助模态集成,利用跨视图注意力或对比目标将异构特征对齐为统一的区域表示,从而提升表示质量。然而,人类移动的时间动态特性仍未得到充分探索。区域流入和流出在一天中波动,区域间连接随时间出现、持续和消失。此外,主流融合策略将视图加性组合,忽略了仅在视图共现时才出现的联合信号。为解决这些不足,我们提出移动流-结构协同(MoSS)方法,从移动数据中派生互补视图:序列视图保留每个区域的小时流入/流出概况,结构视图基于之字形持续图,捕捉区域连接随时间出现、持续和消失的方式。随后,协同模块通过多度交互从这些视图的共现中提取涌现表示,显式捕捉跨视图的高阶信号。在纽约市和芝加哥的大量实验表明,MoSS仅使用移动数据即可在三个下游任务中达到最先进性能,优于依赖辅助模态的基线方法。
英文摘要
Urban region embeddings have shown promising results in diverse urban sensing tasks such as crime, income, and service-call prediction. Recent methods improve representation quality by integrating mobility data with auxiliary modalities, using cross-view attention or contrastive objectives to align heterogeneous features into a unified region representation. However, leveraging the temporal dynamics of human mobility remains under-explored. Regional inflow and outflow fluctuate throughout the day, and inter-region connections emerge, persist, and dissolve over time. Moreover, prevailing fusion strategies combine views additively and miss the joint signal that emerges only when views co-occur. To address these gaps, we propose Mobility Stream-Structure Synergy (MoSS), which derives complementary views from mobility data: a Sequence view that preserves each region's hourly inflow/outflow profile, and a Structure view based on zigzag persistence diagrams that capture how regional connectivity emerges, persists, and dissolves over time. A synergy module then extracts emergent representations from the co-occurrence of these views through multi-degree interactions, explicitly capturing higher-order signal across views. Extensive experiments on New York City and Chicago show that MoSS achieves state-of-the-art performance across three downstream tasks using mobility data alone, outperforming baselines that rely on auxiliary modalities.