arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁在AI审计中做什么?设计用于审计生成式AI的人机协作

Who Does What in AI Auditing? Designing Human-AI Collaboration for Auditing Generative AI

Eunkyu Park, Markelle Roesti, Wesley Hanwen Deng, Renata Barreto, Mohammad Tahaei, Kenneth Holstein, Jason Hong, Motahhare Eslami

arXiv 2609.24986首次发表:更新:

发表机构

Carnegie Mellon University; Seoul National University; eBay; Microsoft Research; EPFL(卡内基梅隆大学; 首尔大学; eBay; 微软研究院; 洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出人机审计协作(HAAC)框架,用于生成式AI审计中的人机分工,通过实验证明AI辅助提升攻击成功率与探索广度,并总结出可操作审计的设计考量。

AI 中文摘要

AI审计日益引入AI智能体以扩大审计覆盖的范围和广度,然而关于审计工作应如何分配而不取代人类判断,目前知之甚少。我们提出了人机审计协作(HAAC),这是一种用于在AI审计中构建人机协作的工作流程和系统。借鉴先前工作并与AI审计从业者进行形成性咨询,HAAC明确了智能体如何在探索、评估、报告和审查方面提供支持,同时在需要情境判断的关键环节保留人类监督。我们将HAAC实例化用于对话式购物智能体,并通过两项研究进行评估。在71名审计员参与的研究中,AI辅助提高了攻击成功率并拓宽了探索范围,同时也影响了后续攻击,并增加了审计员对AI生成评估和报告的依赖。对负责任AI从业者的访谈表明,可操作的审计需要覆盖范围的可见性、可复现的攻击轨迹以及对审计智能体本身的评估。我们的研究结果确定了有效且负责任的人机AI审计的设计考量。

英文摘要

AI auditing increasingly incorporates AI agents to expand the scale and breadth of audit coverage, yet little is known about how auditing work should be divided without displacing human judgment. We introduce Human-Agent Audit Collaboration (HAAC), a workflow and system for structuring human-AI collaboration in AI auditing. Drawing on prior work and formative consultations with AI auditing practitioners, HAAC specifies how agents can support exploration, assessment, reporting, and review while preserving human oversight where contextual judgment is critical. We instantiate HAAC for conversational shopping agents and evaluate it through two studies. With 71 auditors, AI assistance increased attack success and broadened exploration, while also shaping later attacks and increasing auditors' reliance on AI-generated assessments and reports. Interviews with Responsible AI practitioners showed that actionable audits require visibility into coverage, reproducible attack trajectories, and evaluation of the auditing agents themselves. Our findings identify design considerations for effective and accountable human-AI auditing.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑