遵循偏好,错过最优:AI住房推荐中的合规而不优化
Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究审计AI住房推荐,发现模型虽高度合规但39%的推荐被严格支配,存在“合规而不优化”问题,并提出支配率检测作为诊断工具。
AI中文摘要:
大型语言模型正成为消费者在物质利益重大且法律明确领域进行搜索的第一接触点。现有审计表明,模型会根据感知身份引导住房寻找者,但没有人能说明当推荐系统忽略合适选项时用户会损失什么,因为缺乏枚举清单来对遗漏进行评分。我们针对可验证的真实情况对AI住房推荐进行审计。针对纽约市的150个合成租房者场景,我们构建了包含120个真实房源(已知租金、卧室数量和GTFS计算的通勤时间)的候选池,计算满足租房者明确约束的精确集合,并推导其帕累托前沿。主要结果不假设效用函数:如果同一候选池中存在更便宜、通勤时间更短且卧室数量不少于该房源的列表,则该推荐被严格支配。在对来自两家供应商的三个模型的9,945次调用中,合规性近乎完美(违规率1.8%,而随机基线为66.6%),但39.0%的推荐被严格支配,且支配列表的中位价格便宜900美元/月,通勤时间近3.5分钟。场景内操纵区分了通常混淆的两种能力:改变一个句子会使中位推荐租金向正确方向移动646美元/月,因此偏好得到尊重,但推荐仍比同一屏幕上五个最便宜的合格房源高出606美元/月,且在预指定的50美元/月界限的等价性检验下,明确的字典序指令没有带来改进。差距随候选集规模扩大而增大,并在OpenAI和Anthropic模型中重复出现,差异在3美元以内。我们将此失败特征化为“合规而不优化”,提出支配率检测作为可部署的诊断工具,并发布所有代码、提示词和每次调用的结果。
英文摘要:
Large language models are becoming the first point of contact for consumer search in domains where the stakes are material and the law is explicit. Existing audits show that models steer housing seekers by perceived identity, but none can say what a user loses when a recommender overlooks a suitable option, for want of an enumerated inventory to score omissions against. We audit AI housing recommendation against a verifiable ground truth. For each of 150 synthetic renter scenarios in New York City we build a pool of 120 real listings with known rent, bedrooms and GTFS-computed transit commute, compute the exact set satisfying the renter's stated constraints, and derive its Pareto frontier. The primary outcome assumes no utility function: a recommendation is strictly dominated if the same pool holds a listing cheaper, faster to commute from and no smaller in bedrooms. Across 9,945 calls to three models from two vendors, compliance is near-perfect (1.8% violation against a 66.6% random floor), yet 39.0% of recommendations are strictly dominated, and the dominating listing is a median 900 USD/month cheaper and 3.5 minutes closer. A within-scenario manipulation separates two capabilities usually conflated: changing one sentence moves median recommended rent by 646 USD/month in the correct direction, so preferences are honored, yet recommendations still sit 606 USD/month above the five cheapest qualifying listings on the same screen, and an unambiguous lexicographic instruction gives no improvement under equivalence testing against a pre-specified 50 USD/month bound. The gap widens with candidate-set size and replicates across OpenAI and Anthropic models to within 3 USD. We characterize the failure as compliance without optimization, propose dominance-rate instrumentation as a deployable diagnostic, and release all code, prompts and per-call results.