发表机构
Ora Research(Ora Research)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模智能体旅程实验证明,在智能体网络时代,企业网站的可读性(AX)比被外部提及(AEO)更重要,可显著提升推荐率与答案准确率。
AI 中文摘要
2023 年,AI 模型从训练数据中作答,并在数据耗尽时产生幻觉,企业被告知要植入这些知识。此后,模型的训练知识让位于实时网络搜索,相关建议也随之跟进:答案引擎优化(AEO)现在告诉企业在论坛帖子、列表文章和站外引用中散布面包屑,以便 AI 引擎更有可能展示并推荐它们。但仅仅被展示已不再足够:智能体在决定之前会打开结果并阅读它们,而一个买家问题会触发多轮搜索和抓取。在这一深入步骤中,决定结果的是智能体能否抓取并阅读企业自己的网站:智能体体验(AX)。我们认为 AX 是新的 AEO。我们运行了 37,927 次智能体旅程,每次都是关于某个企业的买家问题,跨越四个独立的测试框架,覆盖 1,056 家真实企业,按知名度、先验模型知识和两个 AEO 代理指标进行匹配,然后根据其 AX 水平进行分组。无论网站是否可读,最终答案中仅有 7-10% 来自模型的训练知识。具备智能体就绪(agent-ready)能力的企业,其答案有 78% 的时间基于自身页面构建,而对照组的这一比例为 56%,并且被明确推荐的频率高出 1.9 倍,而针对不具备智能体就绪能力企业的每个有根据的答案,智能体要多花费 64% 的成本。在固定企业、测试框架和问题的情况下,基于网站构建的答案准确率高出 41%。主要的失败并非捏造,而是遗漏:基于网络构建的答案有 3.7 倍的可能性不包含买家所询问的任何事实。四个测试框架之间的基线差异显著,明确推荐率在不同技术栈之间相差七倍,但该效应在每个框架中都成立。在智能体网络时代,可读性胜过被提及,而提升网站的 AX 是企业拥有的最强杠杆。
英文摘要
In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter breadcrumbs across forum threads, listicles, and off-site citations, so AI engines are likelier to surface and recommend them. But being surfaced is no longer enough: an agent opens the results and reads them before deciding, and one buyer question sends it through several rounds of search and fetch. What decides the outcome at this drill-down step is whether the agent can fetch and read the business's own site: agent experience (AX). We argue that AX is the new AEO. We run 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses, matched on fame, prior model knowledge, and two AEO proxies, then split based on their AX level. Only 7-10% of the finished answer comes from the model's training knowledge, whether or not the site is readable. Agent-ready businesses have answers built from their own pages 78% of the time against 56% and are clearly recommended 1.9x more often, while a grounded answer about a not-agent-ready business costs the agent 64% more on average. Holding business, harness, and question fixed, answers built from the site are 41% more accurate on average. The dominant failure is not fabrication but omission: web-built answers are 3.7x more likely to contain none of the facts the buyer asked for. Baselines differ sharply across the four harnesses, with clear-recommendation rates varying sevenfold from stack to stack, yet the recommendation gap holds in every one. In the agentic web era, being readable beats being talked about, and improving a site's AX is the strongest lever a business has.
Comments17 pages, 11 figures