发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM自主Web智能体面临的网页HTML上下文过长问题,提出基于DeBERTa和T5的过滤模型及零样本ColBERT检索器,提升WebArena任务成功率并实现高召回。
AI 中文摘要
由大型语言模型(LLM)驱动的自主Web智能体,因其具备多步推理和决策能力,在自动化各类基于Web的任务中引起了广泛关注。在这些智能体的开发中,一个开放的研究问题在于网页输入的格式。原始HTML源代码包含大量且往往无关的细节,给上下文窗口有限的LLM带来了困难。为应对这一挑战,我们首先在WebArena(Zhou等人,2023)基准上复现了基线模型,如GPT-3.5和LLaMA-2-70B,识别出常见的失败模式。随后,我们提出了两种检索策略,以过滤掉LLM智能体的无关上下文。我们开发了基于DeBERTa和T5的模型,根据元素与任务的相关性对HTML元素进行排序。我们在Mind2Web轨迹数据上对它们进行微调,并将其迁移到WebArena。实验表明,我们的基于DeBERTa的模型将LLaMA-2-70B LLM智能体在WebArena上的成功率从1.97%提升至2.96%。此外,我们开发了一个零样本的基于ColBERT的检索器,其在Mind2Web上的召回率达到0.52,在WebArena上达到0.47。
英文摘要
Autonomous web agents, powered by Large Language Models (LLMs), have garnered significant attention for automating various web-based tasks with multi-step reasoning and decision-making capabilities. An open research question in the development of these agents lies in the format of the webpage input. Raw HTML source code, with its extensive and often irrelevant details, poses difficulties for LLMs with limited context windows. To address this challenge, we first reproduce baseline models such as GPT-3.5 and LLaMA-2-70B on the WebArena (Zhou et al., 2023) benchmark, identifying common failure modes. We then propose two retrieval strategies to filter out irrelevant context for LLM agents. We develop DeBERTa-based and T5-based models that rank HTML elements by their relevance to the task. We fine-tune them on Mind2Web trajectory data and transfer them to WebArena. Experiments show that our DeBERTa-based model improves the success rate of the LLaMA-2-70B LLM agent on WebArena from 1.97% to 2.96%. Moreover, we develop a zero-shot ColBERT-based retriever that is able to retrieve the ground-truth element with a recall of 0.52 on Mind2Web and 0.47 on WebArena.
CommentsAll authors contributed equally to this work and are co-first authors. We conducted the initial research for this paper in 2023