ASSERT:用于生成式人工智能(GenAI)审计的测量流水线
ASSERT: A Measurement Pipeline for GenAI Audits
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究人员提出ASSERT测量流水线,将GenAI审计的报告率与测量选择规范关联,案例显示测量选择会影响报告率及系统排名,可提升审计差异的归因与解释性。
AI中文摘要:
对生成式人工智能(GenAI)系统的审计常将其行为总结为报告率,即被审计系统符合政策的频率。研究人员和利益相关者利用该比率来比较系统、追踪性能退化情况并管控部署。报告率既反映了被审计系统,也体现了其背后的测量选择,因此比率发生变化时,无法明确是系统本身还是测量选择发生了改变。我们提出ASSERT,一种用于GenAI审计的规范驱动型测量流水线,它将每个报告率与用于生成该比率的测量选择的书面规范关联起来。ASSERT可协助起草行为准则和测试用例,随后针对GenAI系统运行审计并返回报告率。在一项关于对话式欺骗的案例研究中,我们观察到报告率会随对话设置、模拟用户、评判者以及不合规证据标准发生显著变化。这些测量选择会大幅改变报告率,甚至可能重新排列GenAI系统的排名。由于每个报告率都与明确的规范相关联,不同审计之间的差异更便于归因和解释。
英文摘要:
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate to compare systems, track regressions, and gate deployment. A reported rate reflects both the system under audit and the measurement choices behind it, so a change in the rate can leave it unclear whether the system or those choices moved. We introduce ASSERT, a specification-driven measurement pipeline for GenAI audits that ties each reported rate to a written specification of the measurement choices used to produce it. ASSERT helps draft a behavioral rubric and test cases, then runs the audit against a GenAI system and returns a reported rate. In a case study on conversational deception, we observe that the reported rate moves substantially with the dialogue setup, the simulated user, the judge, and the evidence bar for non-compliance. These measurement choices substantially change the reported rate and can reorder GenAI system rankings. Because each reported rate is tied to an explicit specification, differences across audits are easier to attribute and interpret.