arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38291cs.CRcs.CL

HARDE:优化智能体执行框架以实现运行时风险检测与执行控制

HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control

Zhuo Liu, Moxin Li, Zhixin Ma, Wentao Shi, Wenjie Wang, Fuli Feng

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体的运行时安全风险,提出风险感知执行框架HARDE,通过两阶段优化整合监控与干预,在三个攻击基准上提升安全性并保持效用。

中文摘要 AI 辅助

大型语言模型(LLM)智能体易受安全风险影响,如注入的恶意指令或误导性信息,这促使需要运行时防御机制,以在多种风险下防止不安全行为执行,同时保持良性任务的效用。现有的系统级防御要么侧重于风险检测而非及时预防,要么依赖预定义规则,在应对多种风险时灵活性有限。我们提出了一种风险感知的执行框架,该框架整合了基于LLM的监控以实现灵活的风险检测,并将监控引导的执行结构化地组织为三个核心模块:触发器、监控器和反馈,从而在限制对良性任务执行干扰的同时实现有针对性的安全干预。为使该框架适应不同的风险和部署场景,我们引入了HARDE,一个两阶段的框架优化流程:首先对每个模块进行独立探测以得出优化指南,然后基于该指南,根据安全性和效用反馈对框架进行迭代优化。在三个攻击基准上的实验表明,HARDE在保持效用的同时提升了运行时安全性,优于手动设计的框架和朴素优化基线。我们的分析表明,有效的运行时防御受益于互补的安全机制、针对攻击的框架优化以及与监控器能力相匹配的框架设计。我们的代码可从此https URL获取。

英文摘要

Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection rather than timely prevention or rely on predefined rules with limited flexibility across diverse risks. We propose a risk-aware harness that integrates LLM-based monitoring for flexible risk detection and structures monitor-guided execution around three core modules: trigger, monitor, and feedback, enabling targeted safety interventions while limiting disruption to benign task execution. To adapt the harness to different risks and deployment settings, we introduce HARDE, a two-stage harness optimization framework that first performs isolated probing of each module to derive an optimization guide, then uses this guide to iteratively optimize the harness based on safety and utility feedback. Experiments across three attack benchmarks show that HARDE improves runtime safety while preserving utility, outperforming manually designed harnesses and naive optimization baselines. Our analysis shows that effective runtime defense benefits from complementary safety mechanisms, attack-aware harness optimization, and harness designs matched to monitor capabilities. Our code is available at https://github.com/Liuz233/HARDE.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • National University of Singapore(新加坡国立大学)
  • Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

↑