AI 中文总结
该研究提出同时评估消费级AI智能体信任与任务完成的方法,在模拟业务环境中测试,发现Fo助手在完成率和信任保持上优于基础模型及开源助手OpenClaw。
AI 中文摘要
行动智能体(Action agents)替人们做事。它们发送电子邮件、花钱、在用户忙于其他事务时给商家打电话,因此一个错误可能在任何人注意到之前就转化为实际行动。它们以两种方式辜负用户:当它们做了用户从未同意的事情,或在用户明确表示继续后却退缩时,它们会破坏信任;当它们放弃那些结果证明难以完成的差事时,它们在任务完成上有所欠缺。这两者都高度依赖于模型周围的“护栏”(harness),即其指令、工具、上下文和防护措施。我们构建了一个评估系统,在模拟世界中,对同一运行过程同时评分信任度和完成度。该模拟世界包含拥有各自网站、收件箱和电话线路的企业,以及会回复消息的人。一个模拟用户会回答助手的问题。信任意味着不会发生用户未同意的事情:没有邮件发送给未经其批准的人,没有私人细节出现在群聊中,没有超出其限额的花费,没有遵循陌生人的指令,也没有在无来源的情况下做出任何声称。每个陷阱都有一个匹配的对照组,在其中采取行动是正确选择。我们利用该评估来衡量Fo(Wajo的个人助理)与三个基础模型上使用基本指令的基础模型,以及与关闭防护措施的Fo护栏的对比。Fo完成了71%的差事,并在94%的陷阱运行中保持了用户的信任。基础模型完成了50%至64%的差事,并在59%至75%的陷阱运行中保持了信任。在匹配的对照组中,Fo稍微较少地继续行动。OpenClaw(一个流行的开源助手,被赋予相同权限)完成了其与Fo共享差事的42%,而Fo为71%,并在共享陷阱运行中保持了74%的用户信任,而Fo为94%。将信任度和完成度放在一起衡量,且针对整个系统而非仅模型本身,我们认为这是使行动智能体能够安全地接管实际工作的方式。
英文摘要
Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold back after the user clearly said go. And they fall short on completion when they give up on errands that turn out to be hard. Both depend heavily on the harness around the model, meaning its instructions, tools, context, and guardrails. We built an evaluation that scores trust and completion on the same runs, in a simulated world of businesses with their own websites, inboxes, and phone lines, and of people who write back. A simulated user answers the assistant's questions. Trust means that nothing happens the user did not agree to. No email goes to someone they never approved, no private detail ends up on a group thread, no money is spent past their limit, no stranger's instructions are followed, and nothing is claimed without a source. Every trap has a matched control in which acting is the right call. We use the evaluation to measure Fo, Wajo's personal assistant, against a base model with basic instructions on three foundation models, and against the Fo harness with its guardrails switched off. Fo completes 71% of the errands and keeps the user's trust on 94% of the trap runs. The base models complete 50% to 64% and keep trust on 59% to 75%. On the matched controls, Fo goes ahead slightly less often. OpenClaw, a popular open-source assistant given the same access, completes 42% of the errands it shares with Fo, against 71%, and keeps the user's trust on 74% of the shared trap runs, against 94%. Measuring trust and completion together, on the whole system rather than the model alone, is how we think action agents become safe to hand real work to.
Comments25 pages, 4 figures, 8 tables