EADC:大语言模型中高级与深层合规性的评估
EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models
浏览论文内容
中文总结 AI 辅助
针对现有LLM合规评估静态基准的局限,本文提出EADC基准,基于AI合规知识图谱与法律专家构建对抗场景,通过4,435+问答对暴露模型监管盲点,保障高级深层合规。
中文摘要 AI 辅助
大语言模型(LLMs)已被广泛应用于各行各业。然而,确保它们遵守复杂的法律和监管框架仍然是一个巨大的挑战。现有的评估范式主要依赖于静态基准,这些基准存在三个严重的局限性:首先,所使用的合规规则不符合人工智能(AI)法律法规的要求;其次,它们仅处理明显的、显式的合规风险,使得隐性和隐蔽的合规风险未被检测到;第三,它们无法沿着逻辑依赖链追踪风险的系统性传播,也无法在基于细微差别的、上下文相关的现实场景中评估合规性。为了弥合这一关键差距,我们引入了EADC,这是一个基于AI合规知识图谱和AI合规法律专家的新型高级评估基准。通过将抽象的法律规则映射为结构化的逻辑多关系图,我们的框架使自动化、不断演化的智能体能够提炼和综合高度复杂的对抗性场景。该合规基准在整个过程中由人类AI法律专家进行审查和修正。由此产生的数据集(4,435多个问答对)提供了一个广泛的多维分类体系,涵盖了关键的监管前沿领域,包括偏见与歧视、公平性、个人隐私保护和价值观。至关重要的是,我们的合规数据集超越了浅层的字符串匹配,通过整合上下文长程交互和逻辑驱动的危险链,捕获了绕过传统过滤器的深层嵌入的合规异常。实验评估表明,我们的框架暴露了最先进LLMs中的关键监管盲点,提供了一个严格的、与AI法律法规一致的基准,以保障LLMs应用中的高级和深层合规性。
英文摘要
Large Language Models (LLMs) have been used in various industries. However, ensuring their compliance with complex laws and regulatory frameworks remains a great challenge. Existing evaluation paradigms mainly rely on static benchmarks that suffer from three severe limitations: First, the compliance rules being used do not comply with the requirements of Artificial Intelligence (AI) laws and regulations; Second, they only handle apparent, explicit compliance risks, leaving implicit and covert compliance risks undetected; Third, they fail to track the systematic propagation of risks along logical dependency chains or evaluate compliance within nuanced, context-based real-world scenarios. To bridge this critical gap, we introduce EADC, a novel advanced evaluation benchmark of LLMs based on an AI compliance knowledge graph and AI compliance legal experts. By mapping abstract legal rules into structured logical multi-relational graphs, our framework enables automated, evolving agents to distill and synthesize highly sophisticated adversarial scenarios. This compliance benchmark is reviewed and corrected by human AI legal experts throughout the whole process. The resulting dataset (4,435+ QA pairs) provides an extensive, multi-dimensional taxonomy covering critical regulatory frontiers, including bias and discrimination, fairness, personal privacy protection, and values. Crucially, our compliance dataset moves beyond shallow string-matching by incorporating contextual long-horizon interactions and logic-driven hazard chains, capturing deeply embedded compliance anomalies that bypass traditional filters. Experiment evaluations demonstrate that our framework exposes critical regulatory blind spots in state-of-the-art LLMs, offering a rigorous, AI laws and regulations-aligned benchmark to safeguard high-level and deep compliance in the application of LLMs.
发表机构
- Tsinghua University(清华大学)
- University of Wisconsin - Madison(威斯康星大学麦迪逊分校)
- University of Chinese Academy of Social Sciences(中国社会科学院大学)
机构由 AI 辅助整理,请以论文原文为准。