arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解本地服务市场中AI提供者推荐

Understanding AI Provider Recommendations in Local Service Markets

Hazem Ibrahim, Yasir Zaki

arXiv 2609.18341首次发表:更新:

发表机构

New York University Abu Dhabi(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计AI在本地服务中的推荐,发现无检索时推荐多为捏造且匹配率低,启用搜索后匹配率显著提升至64-71%,并降低不当行为披露比例和都市规模偏差,表明检索配置对推荐可信度至关重要。

AI 中文摘要

当有人询问AI助手该看哪位医生或该信任哪家机构来管理其储蓄时,答案便是一种推荐。我们对美国100个最大都市区中四个基于注册表的服务领域进行了AI提供者推荐审计,将每项推荐与该领域官方注册表(Medicare临床医生和设施记录,以及SEC顾问披露)进行匹配,并在三种条件下进行:一个开放权重模型、一个无网络搜索的专有模型,以及同一专有模型启用搜索的情况。在没有搜索的情况下,两个模型在网络覆盖稀薄的领域中大多捏造推荐。开放权重模型推荐的医生中仅有4%和专有模型的11%与查询城市中的临床医生匹配,且开放权重模型的匹配属于姓名巧合:其匹配的临床医生成为初级保健医生的可能性并不比从注册表中随机抽取的名字更高。启用搜索后,同一领域中64-71%的推荐与真实提供者匹配。搜索还改变了被推荐的对象。没有搜索时,推荐的咨询公司携带SEC不当行为披露的比例是注册表基准率的3.6倍,即使在调整公司规模后也是如此;启用搜索后,该比例显著低于基准率。在质量和可见性可分别衡量的餐厅领域,显示出3-5倍的评论数量溢价,但评分溢价最多仅为十分之一星。最后,搜索在很大程度上消除了都市规模惩罚:没有搜索时,真实推荐集中在最大的都市区;启用搜索后,各都市规模三分位组的匹配率相似。AI推荐是否可信在很大程度上取决于其检索配置,而非仅取决于底层模型本身,然而,没有检索生成的答案往往不带有其推荐从未被验证的任何迹象。

英文摘要

When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.

Comments12 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑