arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39107cs.AI

MASCRDM:大型语言模型训练过程中合规风险检测与缓解的多智能体系统

MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models

  • Tsinghua University(清华大学)
  • Xi’an Jiaotong-Liverpool University(西交利物浦大学)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • University of Chinese Academy of Social Sciences(中国社会科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

Yan Zhang, Chuming Wei, Ruien Li, Yaoyao Peng, Wusheng Zhang, Guangwen Yang

AI总结:

针对大模型训练中合规风险检测静态且缺乏实时性的问题,提出MASCRDM多智能体系统,基于合规规则与知识图谱,全程监控并给出警报,有效提升合规性且保持语义性能。

AI中文摘要:

大型语言模型(LLMs)已被应用于各个领域。然而,确保LLMs的合规性和安全性,例如避免歧视和偏见,仍然是一个挑战。当前的工作主要集中于检测和过滤训练模型的输入和输出,而非实时研究模型的内在架构。为了应对这一挑战,我们分析了LLMs的训练过程,并发现了两个关键问题:1)现有的大多数方法在检测和过滤方面主要采用静态方式,仅实现局部优化,而没有系统地提升LLMs的合规性。2)现有方法的另一个问题是在整个训练过程中缺乏实时风险检测和缓解,这导致灵活性有限。受此启发,我们提出了MASCRDM(大型语言模型训练过程中合规风险检测与缓解的多智能体系统)。首先,我们基于现有人工智能(AI)法律制定了一套合规规则,并在合规法律专家的指导下构建了一个合规专用的LLM。然后,我们将LLMs分解为多个组件,并基于合规知识图谱识别关键节点。在LLMs训练期间,我们在整个过程中部署多个智能体,为LLM开发者提供合规风险警报和建议。在歧视和偏见基准上的实验表明,我们的多智能体系统能够有效提升合规性,同时保持合理的语义性能。结果表明,我们的方法提供了一条从LLMs内部系统性缓解合规风险的可执行路径。

英文摘要:

Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studying the intrinsic architecture of the models in real-time. To tackle this challenge, we analyze the LLMs training process and discover two critical issues: 1) Most of the existing methods are predominantly static in their approach to detection and filtering, achieving only localized optimizations without systematically enhancing the compliance of LLMs. 2) Another issue with existing approaches is the lack of real-time risk detection and mitigation across the full training process, which leads to limited flexibility. Motivated by these, we propose MASCRDM (Multi-Agent System for Compliance Risk Detection and Mitigation) during the LLM training process. Firstly, we develop a set of compliance rules based on existing Artificial Intelligence (AI) laws and a compliance-specific LLM with the instruction of compliance law experts. Then, we deconstruct LLMs into several components and identify key nodes based on the compliance knowledge graph. During LLMs training, we implement our multiple agents in the whole process, giving compliance risk alerts and suggestions for LLM developers. Experiments on discrimination and bias benchmark demonstrate that our multi-agent system can effectively improve the compliance while maintaining reasonable semantic performance. The results indicate that our method provides an executable path for mitigating compliance risk from within the LLMs systematically.

↑