HxAgent:用于端到端Web应用测试的迭代智能体规划方法
HxAgent: Iterative Agent Planning for End-to-End Web Application Testing
浏览论文内容
中文总结 AI 辅助
HxAgent是一种基于LLM的迭代规划智能体,通过主动修正策略结合多类信息优化Web测试,在MiniWoB++、350项Web任务及OnlineMind2Web数据集上均优于WALT模型,表现优异。
中文摘要 AI 辅助
在自动化Web测试中,利用自然语言功能描述生成测试用例并执行测试,对提升测试效能至关重要,这类任务要求测试智能体在目标应用上自主执行任务并生成测试。我们提出HxAgent,这是一种基于大语言模型(LLM)的迭代规划智能体,具备主动修正策略。每一步操作后,HxAgent会重新评估Web状态,结合三方面信息确定下一步动作:一是当前观测结果,二是过往操作的短期记忆,三是从过往正确/错误操作序列中提取的长期经验。HxAgent在MiniWoB++数据集上实现了97.4%的精确匹配(Exact-Match)准确率,无需人工演示即可达到最优基线水平,且比近期的WALT模型高出10.5%;在包含350个Web任务的数据集上,其精确匹配准确率达83.8%、前缀匹配(Prefix-Match)准确率达91.8%,比WALT模型高出13.4%;在OnlineMind2Web数据集上,它相比WALT模型进一步提升了4.6%。
英文摘要
In automated web testing, generating test cases and performing testing using functionality descriptions in natural-language is crucial for improving efficacy. These tasks require such a testing agent to carry out tasks on the target application and generating tests autonomously. We introduce HxAgent, an iterative LLM-based planning agent with a proactive correction strategy. After each step, HxAgent reassesses the web state to determine the next action using (1) current observations, (2) short-term memory of past actions, and (3) long-term experience extracted from past (in)correct sequences of actions. HxAgent achieves 97.4% Exact-Match accuracy on MiniWoB++, comparable to the best baselines without human demonstrations and surpassing the recent WALT by 10.5%. On a dataset of 350 web tasks, it attains 83.8% Exact-Match and 91.8% Prefix-Match, exceeding WALT by 13.4%. On OnlineMind2Web, it further improves over WALT by 4.6%.