面向高风险机器学习系统的因果证据治理
Causal Evidentiary Governance for High-Risk Machine Learning Systems
- Isik University(伊锡克大学)
- Işık University(伊锡克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对高风险ML系统监管痛点,提出CEG框架,通过DAG划分因果路径、DEP绑定证据,经实证验证其因果归因能力优于传统公平性指标,且具备可行的运算性能。
AI中文摘要:
部署于信贷、招聘及资源分配等场景的机器学习系统正日益受到《欧盟人工智能法案》、《通用数据保护条例》等政策的监管。当前的公平性治理实践依赖观测式公平性指标、事后可解释性及不可变审计日志,但对因果归因与高效证据验证的支持有限。我们提出因果证据治理(CEG)框架,受监管机构需提交版本化有向无环图(DAG),该图将因果路径划分为允许与禁止两类。因果伤害率用于衡量由禁止因果路径导致的预测变异,每项决策附带经签名的决策证据包(DEP),其通过密码学方式将预测与已发布DAG的摘要及路径特定归因绑定,DEP摘要可追加至默克尔树以实现对数级成本的包含性证明。我们采用两层实证方法验证CEG:利用四年PMA信贷监管数据的人口统计摘要,构建四个策略DAG反事实场景下的10000名合成信贷申请人;结果显示,因果伤害率比人口统计 parity 或均等机会更清晰地分离注入的因果效应,跨模型验证与消融研究评估了鲁棒性,对德国信贷数据集的评估表明,关联式公平性指标会大幅低估特定因果路径相关的伤害;最后,概念验证实现展示了操作层面可行的吞吐量,并凸显了相关性能权衡。
英文摘要:
Machine learning systems deployed for credit, hiring, and resource distribution are increasingly subject to regulatory oversight from policies such as the EU AI Act and GDPR. Current fairness governance practices rely on observational fairness metrics, post-hoc explainability, and immutable audit logs, but provide limited support for causal attribution and efficient evidentiary verification. We introduce Causal Evidentiary Governance (CEG), a framework in which regulated institutions commit to a versioned directed acyclic graph (DAG) that partitions causal pathways into allowable and disallowed groups. The Causal Harm Rate measures prediction variation attributable to disallowed causal pathways. Each decision is accompanied by a signed Decision-Evidence Packet (DEP), cryptographically binding the prediction to a digest of the published DAG and path-specific attributions. DEP digests can be appended to a Merkle tree to enable logarithmic-cost inclusion proofs. We validate CEG through a two-layer empirical methodology using demographic summaries from four years of PMA credit supervisory data to construct 10,000 synthetic credit applicants across four strategic DAG counterfactuals. Causal Harm Rate isolates injected causal effects more clearly than demographic parity or equalized odds. Cross-model validation and ablation studies assess robustness. Evaluation on the German Credit dataset shows that harm associated with specific causal pathways can be substantially understated by associational fairness metrics. Finally, a proof-of-concept implementation demonstrates operationally plausible throughput and highlights relevant performance tradeoffs.