arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

terms.txt:面向智能体网络访问的同意与补偿协议

terms.txt: A Consent and Compensation Protocol for Agentic Web Access

Rajarshi Chowdhury

arXiv 2609.11152首次发表:更新:

AI 中文总结

针对AI爬虫破坏开放网络访问约定,提出terms.txt协议,以类似robots.txt的文件定义按路径、按目的的访问条款,并通过签名、令牌和HTTP 402实现同意与补偿,开销极低。

AI 中文摘要

开放网络曾基于一项不成文的约定运行:网站允许爬虫访问,搜索引擎则将访客引导回网站。公开测量显示,这一约定在AI爬虫和智能体面前已被打破。自动化客户端如今占据了大部分请求,训练类爬虫在Cloudflare分类的爬取活动中占主导地位,而最大的AI平台每回引一位访客就会抓取数千个页面。网络的通用控制机制,即robots.txt,无法表达身份、目的、条款或价格,可被绕过,且较新的替代方案大多是专有的CDN功能。我们规范了terms.txt——一种类似robots.txt的文件,用于按路径、按目的定义机器访问条款,并辅以基于源站强制的交换机制,采用Web Bot Auth签名、签名意图、委托令牌、HTTP 402协商和签名收据。我们明确了该交换机制可强制执行、可审计以及可留给合同约定的内容。一个无依赖的实现单次请求在单vCPU上增加0.20至0.65毫秒的开销。

英文摘要

The open web ran on an unwritten bargain: sites admitted crawlers, and search engines sent visitors back. Public measurements show that bargain breaking under AI crawlers and agents. Automated clients now make up most requests, training dominates Cloudflare-classified crawling, and the largest AI platforms fetch thousands of pages for each visitor they return. The web's common control, robots.txt, cannot express identity, purpose, terms, or price, can be circumvented, and newer alternatives are largely proprietary CDN features. We specify terms.txt, a robots.txt-style file for per-path, per-purpose machine-access terms, plus an origin-enforced exchange using Web Bot Auth signatures, signed intent, delegation tokens, HTTP 402 negotiation, and signed receipts. We define what the exchange can enforce, audit, and leave to contract. A dependency-free implementation adds 0.20 to 0.65 ms per request on one vCPU.

Comments7 pages, 1 figure, 2 tables. Submitted to IEEE Internet Computing, Special Issue on Future Internet Systems with LLMs and Agents. Code and raw results: https://doi.org/10.5281/zenodo.22647915

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑