AI治理中的宪法覆盖三难困境
The Constitutional Coverage Trilemma in AI Governance
浏览论文内容
中文总结 AI 辅助
该研究指出前沿AI模型的宪法供给无法覆盖人类需求,提出预算多元主义三难困境,发现精简的宪法菜单可显著降低遗憾,为AI治理提供了关键洞见。
中文摘要 AI 辅助
前沿AI系统发挥着“宪法机构”的作用:每个部署的模型都隐含着安全、有用性、诚实性、自主性和公平性之间的优先级排序。我们研究前沿宪法类型的供给是否覆盖了人类需求。结合对23种前沿大语言模型(LLM)原型的出厂默认宪法的释义控制审计,以及对1649名美国参与者开展的相同工具的成对权衡研究,我们报告三项事实。需求广泛:涵盖所有五种价值,最大的群体占比不到三分之一。供给狭窄且存在漂移:在保守噪声匹配估计下,23种原型的“ hull”( hull指覆盖空间)仅占需求 hull的约2%(在完全审计精度下为0.10%);没有原型将有用性或自主性置于首位(37%的用户在宪法层面无归属);在六个模型家族中,5/6的家族自主性下降,5/6的家族公平性上升,4/6的家族安全性上升,且家族内版本趋势呈单调性(排列置换检验p=0.013),自主性下降集中在安全未受威胁的场景中。漂移的重要性具有方向性:远离已被覆盖不足的价值,机械性地加剧了服务最少用户的福利底线恶化。解决方案的空间有限:二元菜单{e_HON, e_AUT}在平均遗憾上比全部23种原型前沿表现好47%(置信区间[43%,52%]);增加三个顶点可将平均/最差群体遗憾降低多达81%/64%。我们将这些发现形式化为预算多元主义三难困境,证明绑定机制在经验中已实现,并验证结论对基于距离的福利和降级路由具有鲁棒性。该工具和审计框架的详细描述见附录。
英文摘要
Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\sim}2\%$ of the demand hull under conservative noise-matched estimation ($0.10\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \emph{The fix is sparse}: a $2$-vertex menu $\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\}$ beats the full $23$-archetype frontier by $47\%$ on mean regret (CI $[43\%, 52\%]$); three vertex additions cut mean/worst-group regret by up to $81\%$/$64\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.
发表机构
- Kera Health Platforms(Kera健康平台)
机构由 AI 辅助整理,请以论文原文为准。