通过语言模型的视角映射城市
Mapping the City Through the Lens of Language Models
- University College London(伦敦大学学院)
- Centre for Advanced Spatial Analysis (CASA)(高级空间分析中心)
- Tsinghua University(清华大学)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究通过10个开放权重检查点结合多类技术,匿名测量语言模型对城市的隐含假设,勾勒出模型视角下的城市画像,揭示其对城市的偏好倾向。
中文摘要 AI 辅助
语言模型在完成对城市的不明确指代时,往往隐含着关于城市规模、形态、基础设施、环境及功能的未明确假设。我们在不命名具体地点的情况下对这些假设进行了测量,采用10个开放权重检查点,对来自40个经审核指标、7个领域的真实城市形态中心衍生的匿名化概况进行评分。该设计结合了基于概率的约束评分、预先指定的可靠性筛选、谱系感知聚合、多重人口加权、独立重复样本以及全概况验证。最明显的共同倾向是偏好具有更大建成区、近期增长更快、更多已测绘基础设施及非住宅容量、形态更不稀疏的城市概况。大多数符合条件的方向在重复样本中重复出现,且对完整概况的直接评分与按指标构建的评分存在中等程度的一致性。在考虑城市规模和发展后,地理差异缩小,而可靠测量的配对任务表明,典型性与可取性通常密切相关。该框架使模型所认为的“普通城市”这一原本模糊的概念可通过经验追踪,所得证据勾勒出通过语言模型视角呈现的、共享但依赖于模型的城市画像。
英文摘要
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.