arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AURA-Eval:LLM智能体轨迹中风险意识下行动评估框架

AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis, Yu Feng, Aniruddhan Ramesh, Rico Angell, Shang Hong Sim, Chrysoula Zerva, Emmanouil Koukoumidis

arXiv 2609.06783首次发表:更新:

AI 中文总结

AURA-Eval通过受控增强与细粒度诊断,评估LLM智能体在风险意识下的行动安全性,发现无安全路径时模型更易执行不安全行为,前沿专有模型更优。

AI 中文摘要

LLM智能体在操作流程中,不安全的行为可能带来真实后果。现有的安全评估往往将行为简化为单一分数,掩盖了风险识别、行动前检测以及在存在安全解决方案时安全完成任务的能力。我们提出了AURA-Eval,一个将受控增强与工具使用轨迹中行为的细粒度诊断相结合的框架。其流程识别安全关键决策点,生成受控变体,并构建在请求是否存在安全完成路径方面有所不同的对照样本。利用157条来源轨迹,我们生成了1,249个评估项,并评估了20个前沿和开放权重模型。我们制定了评分标准,用于分类风险检测、行动策略以及特定场景下的行动安全性。结果表明,当不存在安全完成路径时,LLM智能体更频繁地从事不安全行为。在这些情况下,前沿专有模型更常识别风险并通过提出替代方案表现出更安全的行为,而被评估的开放权重模型则更常直接执行不安全的请求。在执行前增加影响或减少监督机会也暴露了各模型更大的脆弱性。

英文摘要

LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when a safe solution exists. We introduce AURA-Eval, a framework combining controlled augmentation with granular diagnosis of behavior in tool-use trajectories. Its pipeline identifies safety-critical decision points, generates controlled variations, and constructs counterparts differing in whether a request has a safe fulfillment path. Using 157 sourced trajectories, we generate 1,249 evaluation items and evaluate 20 frontier and open-weight models. We developed rubrics to classify risk detection, action strategy, and scenario-specific action safety. Our results show that LLM agents engage in unsafe behavior more often when no safe fulfillment path exists. In these cases, frontier proprietary models more often recognize risk and exhibit safer behavior by proposing alternatives, while evaluated open-weight models more often directly execute unsafe requests. Increasing impact or reducing opportunities for oversight before execution also exposes greater vulnerability across models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑