arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12056cs.AIcs.HC

为人工智能网络代理设计适用于代理的网站:机器可读性、可操作性和决策可靠性框架

Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

Said Elnaffar, Farzad Rashidi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对人工智能代理介导购物模式下网站设计问题,提出围绕代理可解释性等三个维度的设计框架,通过对照实验评估,结果表明该框架能显著提升人工智能浏览器代理的可靠性与效率。

中文摘要 AI 辅助

在线购物正日益转向一种模式,即人工智能代理为用户独立搜索产品、比较选项、评估约束并执行部分购买流程。网站设计现在必须同时支持人与代理介导的交互。本文介绍了适用于代理的网站,这是一个用于提高电子商务平台对人工智能代理的可读性、可解释性、可验证性和可操作性的设计框架。现有网页设计、SEO和生成引擎优化(GEO)指标不能完全评估网站的代理介导交互能力。所提出的框架围绕代理可解释性、代理可执行性和代理决策可靠性三个维度构建,由机器可读性、语义清晰度、代理可操作性和上下文决策可靠性信号等特征支持。通过一项对照实验对该框架进行评估,比较了一个以人为本的基线和一个相同网站原型的适用于代理的版本,它们具有相同的目录、定价、库存和购物工作流程。评估涉及五项任务、三种浏览器代理模型(GPT - 4.1、Gemini - 2.5 Flash和Grok - 4 Fast)以及300次运行,测量通过、部分通过、失败结果、严格和功能成功率、错误模式、步骤数和令牌消耗。适用于代理的网站在150次运行中有134次通过,而基线为74次(严格成功率分别为89.3%和49.3%),在产品细节提取、比较和多约束选择方面收获最大。它还将部分通过结果从43次减少到3次,并将平均步骤数从9.31降低到6.49。这些结果提供了初步证据,即增强结构清晰度、行动线索、证据信号和时间有效性指标可以显著提高人工智能浏览器代理的可靠性和效率。

英文摘要

Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready website, a design framework for enhancing the readability, interpretability, verifiability, and actionability of e-commerce platforms for AI agents. Existing web design, SEO, and generative engine optimization (GEO) metrics do not fully assess a website's capacity for agent-mediated interaction. The proposed framework is structured around three dimensions agent interpretability, agent executability, and agent decision reliability supported by features such as machine readability, semantic clarity, agent actionability, and contextual decision-reliability signals. The framework is evaluated through a controlled experiment comparing a human-oriented baseline and an agent-ready version of an identical website prototype, with identical catalogs, pricing, stock, and shopping workflows. The evaluation involved five tasks, three browser-agent models (GPT-4.1, Gemini-2.5 Flash, and Grok-4 Fast), and 300 runs, measuring PASS,PARTIAL,FAIL outcomes, strict and functional success rates, error patterns, step counts, and token consumption. The agent-ready website achieved 134 PASS runs out of 150 versus 74 out of 150 for the baseline (strict success rates of 89.3% vs. 49.3%), with the largest gains in product detail extraction, comparison, and multi-constraint selection. It also reduced PARTIAL outcomes from 43 to 3 and lowered the average step count from 9.31 to 6.49. These results provide preliminary evidence that enhanced structural clarity, action cues, evidence signals, and temporal validity indicators can substantially improve the reliability and efficiency of AI browser agents.

发表机构

  • UFR Informatique Université Paris Cité(巴黎西岱大学信息学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑