arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

aiXamine:针对大语言模型安全、安全与隐私的跨维度权衡的统一黑盒评估

aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy

Fatih Deniz, Yazan Boshmaf, Dorde Popovic, Issa Khalil

arXiv 2608.20554首次发表:更新:

发表机构

Qatar Computing Research Institute (QCRI), HBKU(卡塔尔计算研究所(QCRI),哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出aiXamine统一黑盒平台,对超120个LLM开展大规模联合评估,发现安全、隐私等跨维度权衡现象,揭示当前对齐方法将多维度可信性视为单一目标的缺陷。

AI 中文摘要

已部署的大语言模型(LLM)的关键失效模式是跨维度的:某模型在安全对齐上得分99.3,但每三个良性查询中就有一个被弃权(不执行);或在各项能力指标上均有提升,但隐私得分下降21分。现有分别评估安全、安全与隐私的框架无法检测到这些模式。我们推出aiXamine,这是一个统一的黑盒平台,将LLM的可信性作为相互依赖的属性,在安全、安全与隐私维度进行评估。aiXamine通过自动化红队流程协调了9项服务中的46项测试,生成从提示级诊断到跨服务权衡分析的分层风险图谱,可在相同条件下对专有模型和开放权重模型进行可复现的比较。将aiXamine应用于超过120个LLM,完成了5000多次测试运行,我们开展了迄今为止最大规模的安全、安全与隐私联合研究,发现了三个单轴评估无法察觉的跨维度现象。第一,安全执行会产生可量化的安全税:更强的对齐会系统性地增加过度弃权(不执行),迫使提供者在保护与实用性之间做出选择。第二,隐私与其他可信性维度近乎正交,无法被标准对齐所捕捉。第三,我们识别并正式表征了蒸馏诱导的鲁棒性崩溃:无策略内校正的策略外蒸馏会导致熵崩溃,在同一基础架构上鲁棒性从56.9灾难性降至2.6。这些发现,加上规模回报递减和类别依赖的安全行为,表明可信性本质上是多维度的:在一个维度上的进展不能保证甚至会损害其他维度的进展,而当前的对齐方法将其视为单一目标。

英文摘要

The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. Existing evaluation frameworks that assess safety, security, and privacy independently cannot detect these patterns. We introduce aiXamine, a unified black-box platform that evaluates LLM trustworthiness across safety, security, and privacy as interdependent properties. aiXamine orchestrates 46 tests across nine services through an automated red-teaming pipeline, producing hierarchical risk profiles, from prompt-level diagnostics to cross-service trade-off analytics, that enable reproducible comparison of proprietary and open-weight systems under identical conditions. Applying aiXamine to over 120 LLMs through more than 5,000 test runs, we conduct the largest joint safety, security, and privacy study to date and uncover three cross-dimensional phenomena invisible to single-axis evaluation. First, safety enforcement incurs a quantifiable safety tax: stronger alignment systematically increases over-refusal, forcing providers to choose between protection and utility. Second, privacy is near-orthogonal to other trustworthiness dimensions and not captured by standard alignment. Third, we identify and formally characterize distillation-induced robustness collapse: off-policy distillation without on-policy correction causes entropy collapse, catastrophically destroying robustness (56.9$\to$2.6) on the same base architecture. These findings, compounded by diminishing returns from scale and category-dependent safety behaviors, demonstrate that trustworthiness is inherently multi-dimensional: progress along one axis does not guarantee, and can actively undermine, progress along others, yet current alignment methods treat it as a single objective.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑