arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28094cs.AI

一种用于评估治理提案的工具:规模化AI政策分析

An Instrument to Evaluate Governance Proposals: AI Policy Analysis at Scale

Paulo Carvao, Claudio Mayrink Verdun, Isabel Adler, Jeffrey Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种AI政策分析框架,结合专家见解与计算分析,以领域校准模型为基准评估商业LLM,助力用户评估AI治理提案的属性一致性,支持政策制定等人员应对复杂环境。

中文摘要 AI 辅助

本文介绍了一种政策分析框架,用于在不断发展且存在争议的监管环境中系统、透明地评估AI治理提案。AI政策辩论常陷入二元对立立场,掩盖了潜在的权衡与规范性假设。该框架围绕多项政策属性构建政策分析,允许用户明确优先级与张力,无需预设结果。我们采用混合方法,将主题专家的定性见解与计算文本分析相结合,为政策属性 rubric( rubric 可译为评分标准、指标体系)的设计提供依据。这一方法量化了不同政策目标的相对重视程度,并通过比较可视化呈现,支持可解释性与跨政策比较。本文还研究了商业LLM(大型语言模型)在基于rubric的政策分析中的应用,将其输出与具备明确分析假设的领域训练rubric校准模型进行基准测试。该框架不评估政策的有效性或合意性,而是聚焦各属性间的相关性与一致性。通过明确分析假设,包括属性选择、rubric构建及权重方案,该框架使用户能够评估其内置优先级是否与自身规范性承诺一致。该方法不针对特定司法管辖区,旨在支持应对复杂AI治理环境的政策制定者、分析师与研究人员。贡献包括:(1)通过基于实证的rubric进行多维度政策评估,揭示权衡而非解决权衡;(2)透明的混合方法,结合主题专家反馈与计算验证;(3)使用领域训练的rubric校准模型作为基准,比较不同通用大型语言模型。

英文摘要

This paper introduces a policy analysis framework for systematic, transparent assessment of AI governance proposals in an evolving and contested regulatory landscape. AI policy debates often collapse into binary positions that obscure underlying tradeoffs and normative assumptions. The framework structures policy analysis around multiple policy attributes, allowing users to surface priorities and tensions without prescribing outcomes. We use a mixed-methods approach that integrates qualitative insights from subject matter experts with computational text analysis to inform the design of policy attribute rubrics. This quantifies the relative emphasis of different policy objectives and presents them through comparative visualizations that support interpretability and cross-policy comparison. The paper also examines the use of commercial LLMs for rubric-based policy analysis, benchmarking their outputs against a domain-trained rubric-calibrated model with explicitly defined analytical assumptions. Rather than assessing policy effectiveness or desirability, the framework focuses on relevance and alignment across attributes. By making analytical assumptions explicit, including attribute selection, rubric construction, and weighting schemes, the framework enables users to evaluate whether its embedded priorities align with the users' own normative commitments. The approach is jurisdiction-agnostic and intended to support policymakers, analysts, and researchers navigating complex AI governance environments. Contributions: (1) multidimensional policy assessment through empirically grounded rubrics that surface tradeoffs rather than resolving them; (2) a transparent hybrid methodology combining feedback from subject-matter experts with computational validation; and (3) use of domain-trained rubric-calibrated models as a benchmark for comparing different general-purpose large language models.

补充信息

↑