arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAIL Guard:弥合大语言模型智能体负责任人工智能中评估与修复的差距

RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents

Sumit Verma, Pritam Prasun, Pritish Kumar

arXiv 2607.16215首次发表:更新:

发表机构

Responsible AI Labs(负责任人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型智能体评估与修复差距问题,提出RAIL Guard闭环负责任人工智能管道,在多维度评估输出并迭代修复。实验表明其闭环修复收敛率高,能减少不安全执行,还区分了可修复与结构维度问题,系统开源。

AI 中文摘要

现有的大语言模型智能体护栏系统作为二元分类器运行,会阻止不安全内容,这使得组织丢弃失败的输出并从头重试。我们引入了RAIL Guard,这是一个闭环的负责任人工智能管道,它在八个可测量维度上评估大语言模型的输出,并通过评估-重写-重新评估循环迭代修复失败的输出。我们在四个前沿大语言模型和4276个内容输出以及6400个智能体工具调用场景上进行了三个实验来评估该管道。闭环修复的收敛率达到96.9%,而阻止并重试的收敛率为49.1%,不过收敛率最高的方法会使效用降低22.3%;反馈驱动的自我修复在可修复维度上的收敛率达到86.6%,且效用无显著损失(p = 0.177)。工具调用前的评估将不安全的智能体执行减少了33%(p = 0.007),且对任务完成没有影响。我们确定了可响应修复的可修复维度与需要架构而非算法解决方案的结构维度(透明度失败率为93.0%,问责制失败率为92.8%,包容性失败率为82.5%)之间的关键区别。该系统以开源软件开发工具包的形式提供。

英文摘要

Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop responsible AI pipeline that evaluates LLM outputs across eight measurable dimensions and iteratively remediates failing outputs through an evaluate-rewrite-reevaluate loop. We evaluate the pipeline across three experiments on four frontier LLMs and 4,276 content outputs plus 6,400 agent tool-call scenarios. Closed-loop remediation achieves 96.9% convergence versus 49.1% for block-and-retry, though the highest-convergence method reduces utility by 22.3%; feedback-driven self-repair achieves 86.6% convergence on fixable dimensions with no significant utility loss (p = 0.177). Pre-tool-call evaluation reduces unsafe agent executions by 33% (p = 0.007) with zero impact on task completion. We identify a key distinction between fixable dimensions that respond to remediation and structural dimensions (Transparency at 93.0%, Accountability at 92.8%, and Inclusivity at 82.5% failure) that require architectural rather than algorithmic solutions. The system is available as open-source SDKs.

Comments15 pages, 10 figures, 4 tables. Code: https://github.com/Responsible-AI-Labs/rail-score-sdk (Python). Benchmark: https://huggingface.co/datasets/responsible-ai-labs/rail-guard-benchmark

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑