arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05692cs.SE

遗憾主导惊喜:智能体AI安全的设计时需求工程

Regret Dominates Surprise: Design-Time Requirements Engineering for Agentic-AI Safety

Nuwayyir Almohammadi, Rami Bahsoon, Tao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于GORE的MS-RGR机制,利用惊喜和遗憾信号解决智能体安全三难问题,模拟验证可将静默故障降至近零且风险检测快17.5倍,但仅增强而非替代模型级安全训练。

中文摘要 AI 辅助

面向智能体AI领域的需求工程师在评估、规定和实现安全自主性方面面临挑战。主流框架,如面向目标的需求工程(GORE),缺乏在认知不确定性下系统应对这些挑战的机制。我们提出了一种基于GORE的方法,用于对智能体AI系统中的安全自主性进行建模和模拟。我们引入了一种新颖的遗憾主导机制(MS-RGR)来实现安全自主性。MS-RGR使用两种信号:认知惊喜(新颖性检测)和认知遗憾(评估性风险)来解决三难问题:智能体应处于常规自主运行、进行反思推理,还是升级给人类?我们在老年人护理监控和自动驾驶中实例化了MS-RGR。一项100个种子的随机模拟显示,MS-RGR将静默故障降至接近零,并且检测风险的速度比仅使用传感器的基线快约17.5倍,同时通过LTL安全属性保持形式可追溯性。一项回顾性代理实例化,在来自七个LLM的208个AGENTHARM场景的执行轨迹上事后应用DRI门控,表明该门控仅对具有强基线安全性的模型(门控前拒绝率超过80%,例如从84.1%提升至90.9%)改善有害任务拒绝,这表明MS-RGR放大而非替代模型级安全训练。我们讨论了有效性的威胁,将MS-RGR定位为智能体AI需求工程中设计时安全约束的初步可行性证据。

英文摘要

Requirements engineers for agentic-AI domains face challenges in evaluating, specifying, and operationalizing safe autonomy. Mainstream frameworks, such as Goal-Oriented Requirements Engineering (GORE), lack mechanisms to systematically address these challenges under epistemic uncertainty. We contribute an approach that builds on GORE to model and simulate safe autonomy in agentic-AI systems. We introduce a novel Regret-Dominance Mechanism (MS-RGR) to operationalize safe autonomy. MS-RGR uses two signals: epistemic surprise (novelty detection) and cognitive regret (evaluative risk) to address the trilemma problem: should the agent operate in routine autonomy, undergo reflective reasoning, or escalate to human? We instantiate MS-RGR in elderly care monitoring and autonomous driving. A 100-seed stochastic simulation shows MS-RGR reduces silent failures to near-zero and detects risk approximately 17.5 times faster than a sensor-only baseline, remaining formally traceable via LTL safety properties. A retrospective proxy instantiation applying the DRI gate post-hoc over execution traces from 208 AGENTHARM scenarios across seven LLMs shows the gate improves harmful-task refusal only for models with strong baseline safety (over 80% pre-gate refusal, e.g., 84.1% to 90.9%), indicating MS-RGR amplifies rather than substitutes for model-level safety training. We discuss threats to validity, positioning MS-RGR as initial feasibility evidence for design-time safety constraints in agentic-AI requirements engineering.

补充信息

↑