大语言模型能否代表城市公众?经济适用房实验中的行为复制与人口匹配问题
Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment
浏览论文内容
中文总结 AI 辅助
该研究以美国经济适用房调查实验对比8款开放权重LLMs与843名受访者,发现LLMs虽能近似部分总体对比,却无法保留人口结构、组内异质性等关键特征,城市规划的模型评估需测试空间与社会结构的模拟保留情况。
中文摘要 AI 辅助
将大语言模型(LLMs)作为低成本代理来反映城市规划中的居民态度,正受到越来越多的关注。现有研究表明,LLMs可以预测调查实验的平均结果,但对于它们是否保留了这些平均值背后的空间锚定、身份条件结构,即支持度如何随项目临近住宅而变化,以及这种反应如何按产权和党派群体划分,人们知之甚少。我们将8个开放权重LLMs与美国一项经济适用房调查实验中的843名受访者进行比较,测试它们是否能重现当拟建项目从2英里移至1/8英里时,业主与租户之间支持度变化的差异。Qwen 2.5 14B最接近(-0.242,而人类为-0.285),且是唯一符合预设±0.20等效标准的模型;Phi-4 14B方向一致但程度减弱(-0.150),其他模型则表现出较弱、无效或反向的调节作用。这种总体匹配掩盖了结构失效:Qwen减弱了共和党人的对比,夸大了无党派人士的对比,其在27个党派-产权-项目单元格中的均方根误差(RMSE)为0.613,模型与人类的中位数方差比为0.099,问题顺序使对比偏移了+0.367。身份线索去除和选择性非响应改变了可估计的比较,且在20.6%-35.3%的重点比较中,基于理由的响应与匹配的直接选择响应存在差异。因此,LLM可以近似一个总体对比,却无法保留产生该对比的人口结构、组内异质性和测量稳定性。城市规划中的模型评估应测试这种空间和社会结构在模拟中是否得以保留,而非仅测试平均效应。
英文摘要
There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 respondents in a US affordable-housing survey experiment, testing whether they reproduced the owner-renter difference in support change as a proposed development moved from 2 miles to 1/8 mile. Qwen 2.5 14B was closest (-0.242 versus the human -0.285) and was the only model to meet the prespecified +/-0.20 equivalence criterion; Phi-4 14B was directionally aligned but attenuated (-0.150), and other models showed weak, null, or reversed moderation. This aggregate match masked structural failure. Qwen attenuated the Republican contrast and exaggerated the Independent one, its RMSE across 27 party-by-tenure-by-item cells was 0.613, its median model-to-human variance ratio was 0.099, and question order shifted the contrast by +0.367. Identity-cue removal and selective nonresponse changed which comparisons were estimable, and rationale-first responses differed from matched direct-choice responses in 20.6-35.3% of focal comparisons. An LLM can thus approximate one aggregate contrast while failing to preserve the population structure, within-group heterogeneity, and measurement stability that generate it. Model evaluation in urban planning should test whether this spatial and social structure survives simulation, not only average effects.