arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaT$^2$:面向对话智能体黑盒边界测试的自适应测试变换

AdaT$^2$: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents

Liting Lin, Boxi Yu, Qinghua Xu, Yuzhong Zhang, Lionel Briand, Emir Muñoz

arXiv 2610.10141首次发表:更新:

发表机构

Lero, the Research Ireland Centre for Software, University of Limerick; The Chinese University of Hong Kong, Shenzhen; University of Ottawa; Genesys(利默里克大学Lero爱尔兰软件研究中心; 香港中文大学(深圳); 渥太华大学; Genesys公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AdaT$^2$方法,利用LLM提取陈述并自适应选择变换指令,生成有效边界测试,在四个领域上优于AgentEval,并能检测更多策略故障。

AI 中文摘要

基于大语言模型(LLM)的对话智能体必须遵守策略。策略中的每个条件在用户请求之间划定一条边界,智能体在边界两侧必须表现出不同的行为。我们提出了AdaT$^2$,它从与充当用户的LLM进行的探索性对话中提取智能体回复中的陈述,并利用这些陈述来指导边界测试生成。每条陈述描述一个条件以及在该条件成立时期望的行为。除了由单一陈述引导的普通测试外,AdaT$^2$还编写由陈述和测试变换指令(例如“省略一个必需输入”)组成的配对引导的变换测试。配对中的指令可以将测试移动到陈述边界的另一侧或移动到另一条边界。陈述和指令形成的配对数量远多于一次运行所能尝试的数量,且许多配对不适用。因此,自适应配对选择根据新颖性选择每条配对的陈述,并使用bandit算法Bayes-UCB选择指令,该算法根据先前的配对是否产生了测试以及智能体是否根据LLM评判者通过了测试来学习。我们的基准分别统计每条边界的两侧,并区分由智能体的提示、工具代码或知识库明确定义的边界与智能体的LLM从领域知识推断出的边界。在$\ au^3$-bench的四个领域上,AdaT$^2$的测试中62.7%至83.3%是有效的边界测试,其预期行为被明确定义,在每个领域都高于AgentEval的47.7%至68.1%。变换测试增加了13至46条普通测试遗漏的明确定义边界。作为回归测试,AdaT$^2$的测试套件在航空领域检测到全部八个种子策略故障,在零售领域检测到八个中的四个,而AgentEval的测试套件(测试数量不到其三分之一)分别检测到五个和两个。

英文摘要

Conversational agents based on large language models (LLMs) must comply with policies. Each condition in a policy draws a boundary between user requests, and the agent must behave differently on its two sides. We present AdaT$^2$, which extracts statements from the agent's replies in exploratory conversations with an LLM acting as the user, and uses the statements to guide boundary test generation. Each statement describes one condition and the behavior expected when the condition holds. Besides plain tests guided by single statements, AdaT$^2$ writes transformed tests guided by pairs of a statement and a test transformation instruction, such as "omit one required input". The instruction of a pair can move a test to the other side of the statement's boundary or to another boundary. The statements and instructions form far more pairs than a run can try, and many pairs are not applicable. Adaptive pair selection therefore chooses the statement of each pair by novelty and the instruction with the bandit algorithm Bayes-UCB, which learns from whether earlier pairs yielded a test and whether the agent passed it according to an LLM judge. Our benchmark counts the two sides of each boundary separately and distinguishes boundaries explicitly defined by the agent's prompt, tool code, or knowledge base from boundaries that the agent's LLM infers from domain knowledge. On four domains of $τ^3$-bench, 62.7% to 83.3% of AdaT$^2$'s tests are valid boundary tests whose expected behavior is explicitly defined, higher in every domain than AgentEval's (47.7% to 68.1%). Transformed tests add 13 to 46 explicitly defined boundaries that plain tests miss. As regression tests, AdaT$^2$'s test suites detect all eight seeded policy faults in the airline domain and four of eight in the retail domain, and AgentEval's test suites, with fewer than a third as many tests, detect five and two, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑