arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21413cs.CE

合成人类移动数据生成:表示、方法与实际能力的结构化综述

Synthetic Human Mobility Data Generation: A Structured Review of Representations, Methods, and Practical Capabilities

  • Centre for Advanced Spatial Analysis, University College London(伦敦大学学院高级空间分析中心)
  • Geospatial Data Science Lab, Department of Geography, University of Wisconsin-Madison(威斯康星大学麦迪逊分校地理系地理数据科学实验室)
  • Center for Spatial Information Science, the University of Tokyo(东京大学空间信息科学中心)

机构由 AI 辅助整理,请以论文原文为准。

Yanbo Pang, Chen Zhong, Song Gao, Yoshihide Sekimoto

AI总结:

本文从城市分析视角结构化综述合成人类移动数据生成方法,提出意义-人口-自主性框架,指出多数方法难以同时满足行为意义、人口基础和自主生成需求。

AI中文摘要:

人类移动数据已成为城市分析中日益重要的组成部分。尽管可用的移动数据源范围已大幅扩展,但其获取仍受到商业限制、隐私问题和制度壁垒的严重约束。数据保护程序也常常降低已发布数据集的分析价值。合成移动数据已成为一种有前景的解决方案,但现有方法在底层机制、所保留的信息、生成的输出以及能够支持的分析问题方面存在显著差异。它们在城市分析中的相对优势和权衡尚未得到充分理解。本文从城市分析视角对合成人类移动数据生成进行了结构化综述。我们按方法论家族对文献进行梳理,并根据每个家族原生生成的移动输出以及这些输出所支持的分析能力对其进行索引。我们首先提供了合成数据产品的分类法,包括人口与个体表示、活动日程、出行与行程记录、轨迹以及聚合移动模式。随后,我们回顾了主要的方法论家族,涵盖机制模型、基于调查的人口合成、基于活动和基于智能体的模拟、深度生成模型、基于Transformer的移动语言模型以及LLM智能体系统。基于这一综合,我们引入了一个意义-人口-自主性框架,该框架从三个维度刻画这些方法:行为意义、人口基础与规模以及生成自主性。我们认为这些维度是下游城市分析的主要需求。很少有方法能同时提供行为意义、人口基础和自主生成,而能在生成中约束于可行轨迹的方法则更少。

英文摘要:

Human mobility data has become an increasingly important component of urban analytics. Although the range of available mobility data sources has expanded substantially, access remains highly constrained by commercial restrictions, privacy concerns, and institutional barriers. Data protection procedures also often reduce the analytical value of released datasets. Synthetic mobility data has emerged as a promising solution, but existing methods differ substantially in their underlying mechanisms, the information they preserve, the outputs they generate, and the analytical questions they can support. Their comparative strengths and trade-offs remain insufficiently understood for urban analytics. This paper presents a structured review of synthetic human mobility data generation from an urban analytics perspective. We review the literature by methodological family and index it by the mobility outputs each family generates natively and the analytical capabilities those outputs enable. We first provide a taxonomy of synthetic data products, including population and persona representations, activity schedules, trip and tour records, trajectories, and aggregate mobility patterns. We then review the major methodological families, spanning mechanistic models, survey-driven population synthesis, activity- and agent-based simulation, deep generative models, transformer-based mobility language models, and LLM-agentic systems. Building on this synthesis, we introduce a Meaning-Population-Autonomy framework that characterises these methods along three dimensions: behavioural meaning, population grounding and scale, and generation autonomy. We consider these dimensions the principal requirements for downstream urban analytics. Few methods deliver behavioural meaning, population grounding and autonomous generation at once, and fewer still with generation constrained to feasible trajectories.

补充信息

↑