金融智能体的自我进化审计:能力提升、安全漂移与执行接口不匹配
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch
浏览论文内容
中文总结 AI 辅助
该研究通过模拟电子银行环境审计SkillOpt、AWM、ReasoningBank等自我进化金融智能体,发现仅靠准确率评估不足,需跟踪退化、攻击面接触等指标,且AWM存在执行接口不匹配的评估风险。
中文摘要 AI 辅助
自我进化智能体将经验转化为可复用的技能、工作流或记忆,但仅进化后的准确率无法表明学习到的行为是否保留了先前的正确行为或安全性。我们在模拟电子银行环境中,使用匹配的良性获取轨迹、密封的评估端点、基于执行的检查以及独立的状态回放,对SkillOpt、Agent Workflow Memory(AWM,智能体工作流记忆)和ReasoningBank进行审计。在Qwen 3.7 Flash模型上,SkillOpt将良性效用从0.741提升至0.837,而对注入内容的暴露度从0.820上升至0.943;暴露后的条件攻击成功率从0.605降至0.562,但整体攻击成功率(ASR)从0.496升至0.530,未授权的金融状态变化升至0.685。在三个独立进化的谱系中,能力、暴露度和未授权状态变化在所有三个谱系中均增加,而ASR仅在两个谱系中上升。ReasoningBank将效用提升至0.859,且未增加总体ASR,尽管未授权状态变化仍略高于静态模型。AWM揭示了另一种评估风险:字面的WebArena文本动作信封会在我们的原生函数调用执行器中扰乱工具执行;在事后敏感性测试中,仅移除该信封就将效用从0.319恢复至0.756,同时暴露度从0.299升至0.909,ASR从0.195升至0.575。因此,审计自我进化的金融智能体需要跟踪退化、攻击面接触、未授权金融状态变化以及工件-执行器兼容性,而非仅依赖准确率。
英文摘要
Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution accuracy alone does not show whether learned behavior preserves previously correct behavior or security. We audit SkillOpt, Agent Workflow Memory (AWM), and ReasoningBank in simulated e-banking using matched benign acquisition trajectories, sealed evaluation endpoints, execution-grounded checks, and independent state replay. On Qwen 3.7 Flash, SkillOpt raises benign utility from 0.741 to 0.837 while exposure to injected content rises from 0.820 to 0.943. Conditional attack success after exposure falls from 0.605 to 0.562, yet overall attack success rate (ASR) rises from 0.496 to 0.530 and unauthorized financial state changes rise to 0.685. Across three independently evolved lineages, capability, exposure, and unauthorized-state changes increase in all three, whereas ASR increases in only two. ReasoningBank raises utility to 0.859 without increasing aggregate ASR, although unauthorized state changes remain slightly above Static. AWM reveals a separate evaluation hazard: a literal WebArena text-action envelope disrupts tool execution in our native function-calling executor. In a post-hoc sensitivity test, removing only that envelope restores utility from 0.319 to 0.756, while exposure rises from 0.299 to 0.909 and ASR from 0.195 to 0.575. Auditing self-evolving financial agents therefore requires tracking regressions, attack-surface contact, unauthorized financial-state change, and artifact-executor compatibility, not accuracy alone.