AI 中文总结
研究探讨生成式AI助手是否遵守robots.txt,对十个有网络搜索功能的AI助手进行实证研究,通过特定配置和条件评估其对http URL规则的遵守情况,发现不同助手差异大,结果引发相关法律治理问题及对更新标准的需求。
AI 中文摘要
人工智能助手在推理时越来越多地检索网络内容以提供新颖且有依据的答案,但尚不清楚这些搜索增强功能是否尊重通过http URL表达的网站所有者限制。我们对十个具有网络搜索功能的广泛使用的人工智能助手进行了实证研究。首先确定产生可观察网络浏览行为的配置并记录检索期间暴露的用户代理。然后在四个互补条件下评估对http URL规则的遵守情况。通过服务器端日志和嵌入目标页面的密码,区分实际页面访问和用户可见答案的正确性。结果显示不同助手存在很大差异,引发了关于人工智能辅助网络访问是否充分尊重内容所有者权利和限制的法律和治理问题,也凸显了更新、可执行标准的迫切需求。
英文摘要
AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots$.$txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed during retrieval. We then evaluate compliance with controlled robots$.$txt rules across four complementary conditions: allowed for all user-agents, disallowed for all user-agents, allowed only for the assistant-specific user-agent, and disallowed only for that user-agent. Using server-side logs and secret codes embedded in target pages, we distinguish actual page access from user-visible answer correctness across 200 trials. Our results show substantial variation across assistants. Some systems followed the expected allowed/disallowed access pattern, whereas others accessed restricted resources without requesting robots$.$txt or used generic user-agents that complicated attribution. We also find that retrieval behavior and answer correctness can diverge: assistants may access pages without surfacing the retrieved content, or fail to access even allowed resources. These findings raise broader legal and governance concerns about whether AI-assisted web access adequately respects content owners' rights and restrictions. Furthermore, our observations provide valuable insight into the growing erosion of traditional web governance protocols, highlighting the urgent need for updated, enforceable standards that guarantee publisher autonomy in the age of search-augmented AI assistants.