arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08082cs.AIcs.CL

POLAR:面向工具调用LLM智能体的本体引导风险预防

POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

Yunju Kang, Seonghyeon Cho, Irene Li, Yeo-Chan Yoon, Chanjun Park

首次发表
浏览论文内容

中文总结 AI 辅助

POLAR提出基于两层本体的护栏框架,通过可逆性评分剪除高风险工具调用,在τ²基准上部分提升任务奖励,但存在领域与模型间的效用权衡。

中文摘要 AI 辅助

LLM工具使用智能体在动态环境中运行,其中许多操作带有操作风险。然而,大多数安全机制仅在错误显现后才做出反应。现有的预防性方法要么在思维链推理上对智能体进行微调,要么将自然语言护栏编译为运行时检查,但这些方法都没有提供结构化的、可审计的判定结果。我们提出POLAR,一个面向小型工具调用智能体的护栏框架,通过一个结构化的两层本体评估可逆性。POLAR通过推导候选逆序列为每个操作分配一个分级可逆性分数;未通过阈值的调用在执行前被剪除。在τ²基准上对六个智能体模型进行评估,POLAR在航空领域将四个智能体中的六个平均任务奖励提高了0.11至0.18分,但十八个模型-领域组合中仅有八个整体改善;零售领域和更强的智能体往往出现回退。POLAR提供了可审计的结构化检查,并刻画了其任务效用的权衡。奖励并非预防伤害的直接衡量标准。

英文摘要

LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on $τ^2$-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.

发表机构

  • Soongsil University(崇实大学)
  • Tokyo University(东京大学)
  • Jeju National University(济州国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑