迷失在请求中:通信变异如何干扰邮件代理中的检索与行动
Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents
浏览论文内容
中文总结 AI 辅助
本研究测试电子邮件助手在请求表达方式变化(如间接、正式或方言)下的鲁棒性,发现这些变化会降低性能,并强调评估应变化请求表达并单独测量任务完成情况。
中文摘要 AI 辅助
电子邮件助手不应仅仅因为用户以不同方式表达相同请求而完成更少的工作。然而,大多数基准测试仅使用一个规范请求来测试每个任务,导致这种鲁棒性在很大程度上未被测量。我们测试了当请求信息、可用证据和预期结果保持不变,但通信风格或英语变体发生变化时,电子邮件助手是否仍然可靠。我们构建了沿五个通信风格轴和四个基于规则方言条件的验证变体,并在三个基准上评估它们:一个检索增强生成(RAG)流水线和两个工具使用代理。间接请求降低了所有三个基准的性能,而正式请求降低了两个代理基准的性能。更仔细地检查系统表明,这些失败有不同的原因。冗长请求主要通过使相关电子邮件更难找到来损害词汇检索器。相比之下,即使检索到相关电子邮件,间接和方言变体仍然有害。在代理设置中,间接和正式请求主要导致代理省略所需行动,而不是采取更多无根据的行动。这些结果表明,成功的响应不足以建立鲁棒性:评估应变化请求的表达方式,并单独测量代理是否完成请求的工作。
英文摘要
An email assistant should not complete less work simply because a user phrases the same request differently. Yet most benchmarks test each task with only one canonical request, leaving this form of robustness largely unmeasured. We test whether email assistants remain reliable when the requested information, available evidence, and expected outcome stay fixed, but the communication style or English variety changes. We construct validated variants along five communication-style axes and four rule-based dialect conditions, and evaluate them on three benchmarks: a retrieval-augmented generation (RAG) pipeline and two tool-using agents. Indirect requests reduce performance on all three benchmarks, while formal requests reduce performance on both agentic benchmarks. Examining the systems more closely shows that these failures have different causes. Verbose requests mainly hurt a lexical retriever by making the relevant email harder to find. By contrast, indirect and dialect variants remain harmful even when the relevant email is retrieved. In the agentic setting, indirect and formal requests mainly cause the agents to omit required actions, not to take more unsupported actions. These results show that a successful response is not enough to establish robustness: evaluations should vary how requests are expressed and separately measure whether agents complete the requested work.