AI 中文总结
本研究对巴厘岛两个区域的4776家餐饮场所开展完整普查,审计四大AI助手的推荐结果,发现多数场所未被推荐、存在过时推荐问题,且跨系统一致性低。
AI 中文摘要
AI助手正成为本地探索的主要界面,但人们几乎不了解它们推荐哪些场所——尤其是在餐饮领域,推荐会直接影响营收。我们开展了首次以普查为基准的AI场所推荐审计:完整枚举了巴厘岛仓古和乌布两个限定市场的4776家咖啡馆、餐厅和酒吧,据此评估四个生产级AI系统(ChatGPT、Claude、Gemini、Perplexity)针对96个人设条件查询的2208个搜索响应,这些数据是按预注册协议在7天内收集的。由于我们观测了完整市场,得以测量抽样审计无法覆盖的情况:85.6%的场所从未被任何系统推荐,其中72.6%甚至是拥有50条及以上评分的成熟场所。可见性遵循双边际结构:进入答案与文档相关,包括评论数量(优势比OR=1.64)、拥有自有网站(OR=1.92)、列出价格信息(OR=1.54)以及第三方网页提及(OR=1.44);而星级评分在此边际无关联(OR=0.89)。答案内的排名则反转了该模式:在被推荐的场所中,评分显著预测首位位置(OR=1.17)。在开放POI数据集(Foursquare)中的存在性,这一民间理论的可见性因素,在两个边际均无正向影响。完全编造的情况极少(提及的0.08%),但系统93次推荐了永久关闭的场所——过时而非幻觉是实际失败模式。跨系统一致性较低(前20名Jaccard指数为0.33至0.54)。为期两周的重测显示,跨时段答案相似性与当日重复运行的相似性相当:变动源于抽样随机性,而非时间漂移。我们公开了协议、注册构建方法及衍生数据。
英文摘要
AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.
Comments31 pages, 10 figures