发表机构
Shanghai Artificial Intelligence Laboratory; Harbin Institute of Technology; Fudan University; Zhejiang University; Hefei University of Technology(上海人工智能实验室; 哈尔滨工业大学; 复旦大学; 浙江大学; 合肥工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对计算机使用智能体仅优化任务完成而缺乏安全行为的问题,提出SCOPE联合后训练方法及SCOPE-Gen数据流水线,通过监督微调和在线强化学习,在OSWorld和OS-BLIND上取得最佳能力-安全平衡。
AI 中文摘要
计算机使用智能体(CUA)在通过图形用户界面完成复杂任务方面取得了快速进展,然而仅以任务成功为中心的后期训练并不能引发可靠的安全行为。一个可靠的CUA必须根据风险来调节其执行:它应完成普通的良性任务,避免环境危害并在存在安全完成路径时继续执行,而当目标有害或不存在安全路径时则拒绝执行。为了学习这种条件策略,我们开发了策略执行的安全与能力优化(SCOPE)方法,该方法联合后训练一个CUA,使其同时具备任务执行能力和安全感知决策能力。为了为这一联合目标提供对齐的训练数据,我们进一步引入了SCOPE-Gen,一个自动化流水线,用于合成可验证的能力任务并将其转换为配对环境风险变体,同时保留其原始目标。利用生成的任务,我们构建了SATraj-OS,一个包含能力演示、安全继续和明确拒绝的轨迹数据集。SCOPE首先通过监督微调学习所有三种轨迹类型,然后通过在线强化学习进一步提高任务完成率。从Qwen3.5-9B开始,SCOPE-RL在OSWorld上实现了54.17%的任务成功率,在OS-BLIND上实现了64.30%的攻击规避率,在被评估的智能体中产生了最佳的聚合能力-安全得分58.80%。消融实验揭示了两种安全监督形式的不对称但互补的作用:拒绝轨迹贡献了大部分攻击规避增益,而风险处理轨迹在可比的攻击规避水平下保留了更大的任务效用。
英文摘要
Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete ordinary benign tasks, avoid environmental hazards and continue when a safe completion path remains, and refuse when the goal is harmful or no safe path exists. To learn this conditional policy, we develop Safety and Capability Optimization for Policy Execution (SCOPE), which jointly post-trains a CUA for task-execution capability and safety-aware decision making. To provide aligned training data for this joint objective, we further introduce SCOPE-Gen, an automated pipeline that synthesizes verifiable capability tasks and converts them into paired environment-risk variants while preserving their original goals. Using the resulting tasks, we construct SATraj-OS, a trajectory dataset comprising capability demonstrations, safe continuations, and explicit refusals. SCOPE first learns from all three trajectory types through supervised fine-tuning and then further improves task completion through online reinforcement learning. Starting from Qwen3.5-9B, SCOPE-RL achieves a 54.17% task success rate on OSWorld and a 64.30% attack-avoidance rate on OS-BLIND, yielding the best aggregate capability--safety score of 58.80% among the evaluated agents. Ablations reveal asymmetric but complementary roles for the two forms of safety supervision: refusal trajectories account for most of the attack-avoidance gain, whereas risk-handling trajectories preserve greater task utility at comparable attack-avoidance levels.
CommentsCorrected an author name typo in the metadata; manuscript unchanged