arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07446cs.SEcs.AIcs.CY

分类驱动的开源AI风险缓解工具分析

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

Afreen Alam, Evgenija Popchanovska, Ana Gjorgjevikj, Maryan Rizinski, Lubomir T. Chitkushev, Irena Vodenska, Dimitar Trajanov

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出分类驱动的协议,将21种开源AI风险缓解工具映射到MIT分类法,发现工具在治理等领域存在缺口,为企业AI风险缓解提供实用框架。

中文摘要 AI 辅助

大型语言模型(LLMs)在企业环境中的快速应用带来了运营、安全和治理风险。随着生成式AI应用从试点转向生产,手动危害识别与缓解难以规模化。尽管许多工具支持模型评估、对抗测试、运行时防护和可观测性,但工具生态仍碎片化。这些工具通常针对特定工程任务设计,且所用技术术语与治理框架或风险分类法不匹配,导致难以确定哪些工具对应哪些风险,以及关键缺口所在。本文提出一种结构化协议,通过对开源LLM评估与安全工具进行分类驱动分析,实现AI风险缓解自动化。我们将21种主流开源工具的能力映射到扩展MIT AI风险缓解与响应分类法的32个子类别。一种LLM辅助的检索增强生成(RAG)流水线分析源代码和文档,以提取每个分类类别的能力。可靠性评估显示,三名独立评审者间存在中等一致性(Fleiss' Kappa = 0.509)。分析揭示工具生态高度失衡:工具集中在技术和运营控制方面,而治理、法律与监管、财务与市场控制基本未被覆盖。这促使构建分层风险缓解架构,将基于工具的控制与组织和监管流程相结合。多数投票后,该映射协议的F1分数达75.5%。总体而言,本研究提供了企业AI风险类别与开源缓解能力间的实用映射,明确了仍需人工监督的环节,并提出了适用于开源及专有解决方案的分类驱动框架。

英文摘要

Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI applications move from pilot to production, manual harm identification and mitigation are becoming difficult to scale. Although many tools support model evaluation, adversarial testing, runtime guardrails, and observability, the tooling landscape remains fragmented. Tools are typically designed for specific engineering tasks and described in technical terms that do not align with governance frameworks or risk taxonomies, making it difficult to determine which tools address which risks and where critical gaps remain. This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools. We map the capabilities of 21 prominent open-source tools to the 32 subcategories of the extended MIT AI Risk Mitigation and Response Taxonomy. An LLM-assisted retrieval-augmented generation pipeline analyzes source code and documentation to extract capabilities for each taxonomy category. Reliability assessment yielded moderate agreement (Fleiss' Kappa = 0.509) among three independent reviewers. The analysis reveals a highly skewed landscape in which tools cluster around technical and operational controls, while governance, legal and regulatory, and financial and market controls remain largely unaddressed. This motivates a layered risk-mitigation architecture combining tool-based controls with organizational and regulatory processes. The mapping protocol achieved an F1 score of 75.5% after majority voting. Overall, the study provides a practical mapping between enterprise AI risk categories and open-source mitigation capabilities, identifies where human oversight remains necessary, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.

发表机构

  • Boston University(波士顿大学)
  • Ss. Cyril and Methodius University(圣西里尔与美多德大学)

机构由 AI 辅助整理,请以论文原文为准。

↑