arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式人工智能助手是否遵守robots.txt?追踪可见答案之外的网络访问

Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers

Gabriel Lopez-Fonseca, David Rodriguez, Stefan Bechtold, Jose M. Del Alamo

arXiv 2607.14447首次发表:更新:

AI 中文总结

研究探讨生成式AI助手是否遵守robots.txt,对十个有网络搜索功能的AI助手进行实证研究,通过特定配置和条件评估其对http URL规则的遵守情况,发现不同助手差异大,结果引发相关法律治理问题及对更新标准的需求。

AI 中文摘要

人工智能助手在推理时越来越多地检索网络内容以提供新颖且有依据的答案,但尚不清楚这些搜索增强功能是否尊重通过http URL表达的网站所有者限制。我们对十个具有网络搜索功能的广泛使用的人工智能助手进行了实证研究。首先确定产生可观察网络浏览行为的配置并记录检索期间暴露的用户代理。然后在四个互补条件下评估对http URL规则的遵守情况。通过服务器端日志和嵌入目标页面的密码,区分实际页面访问和用户可见答案的正确性。结果显示不同助手存在很大差异,引发了关于人工智能辅助网络访问是否充分尊重内容所有者权利和限制的法律和治理问题,也凸显了更新、可执行标准的迫切需求。

英文摘要

AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots$.$txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed during retrieval. We then evaluate compliance with controlled robots$.$txt rules across four complementary conditions: allowed for all user-agents, disallowed for all user-agents, allowed only for the assistant-specific user-agent, and disallowed only for that user-agent. Using server-side logs and secret codes embedded in target pages, we distinguish actual page access from user-visible answer correctness across 200 trials. Our results show substantial variation across assistants. Some systems followed the expected allowed/disallowed access pattern, whereas others accessed restricted resources without requesting robots$.$txt or used generic user-agents that complicated attribution. We also find that retrieval behavior and answer correctness can diverge: assistants may access pages without surfacing the retrieved content, or fail to access even allowed resources. These findings raise broader legal and governance concerns about whether AI-assisted web access adequately respects content owners' rights and restrictions. Furthermore, our observations provide valuable insight into the growing erosion of traditional web governance protocols, highlighting the urgent need for updated, enforceable standards that guarantee publisher autonomy in the age of search-augmented AI assistants.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑