GoT-CD:思维图因果发现与事后路径特定公平性审计的脆弱性
GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits
浏览论文内容
中文总结 AI 辅助
本文提出GoT-CD方法,其生成的DAG在因果发现基准中表现优异,但研究发现事后路径特定公平性审计对因果发现的结构误差较为脆弱,需结合结构发现开展路径特定公平性分析。
中文摘要 AI 辅助
因果发现可从观测数据中恢复有向结构,且正越来越多地应用于临床环境,以支持机制推理和预测模型的公平性审计。路径特定反事实公平性关注受保护属性是否通过非合法路径影响结果,但这些估计量是相对于给定因果图定义的,因此会继承发现步骤引入的任何误差。发现方法通常以同等权重所有边的聚合结构指标进行评估,且没有既定评估会询问审计所依赖的特定路径是否在发现过程中保留,或当该路径缺失时审计会报告什么。本文表明,全图思维图(Graph-of-Thoughts)推理可生成与大语言模型(LLM)基线在结构上具有竞争力的无环发现图,但仅结构保真度无法保证公平性忠实的审计。我们提出GoT-CD,其中推理单元为完整候选边集:并行生成多个图,通过确定性有效性函数评分,并在禁止虚构边的硬并约束下合并,同时在提交前通过贪心投影强制生成有向无环图(DAG)。GoT-CD在所有五个报告的基准测试中均返回有效DAG,且在Asia、Alzheimer's和COVID-Respiratory数据集上,其DAG-有效F1得分在LLM方法中达到最佳。在具有已知不公平路径的Alzheimer's基准测试中,事后路径特定审计显示,八个发现图中有五个未恢复从敏感属性到结果的路径,因此报告整体效应为零,而中介效应仍然存在,这需要在结构发现之外开展下游路径特定公平性分析。
英文摘要
Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. Path-specific counterfactual fairness asks whether a protected attribute influences an outcome through illegitimate pathways, but these estimands are defined relative to a supplied causal graph and therefore inherit whatever errors the discovery step introduces. Discovery methods are routinely scored on aggregate structural metrics that weight all edges equally, and no established evaluation asks whether the specific pathway an audit depends on survives discovery---or what the audit reports when that pathway is missing. Here we show that full-graph Graph-of-Thoughts reasoning yields acyclic discovered graphs that are structurally competitive with large language model (LLM) baselines, yet that structural fidelity alone does not guarantee fairness-faithful audits. We introduce GoT-CD, in which the reasoning unit is a complete candidate edge set: multiple graphs are generated in parallel, scored by a deterministic validity function, and merged under a hard union constraint that forbids invented edges, with greedy projection enforcing a DAG before commitment. GoT-CD returns a valid DAG on all five reported benchmarks and achieves the best DAG-valid F1 score among LLM methods on Asia, Alzheimer's, and COVID-Respiratory datasets. On an Alzheimer's benchmark with known unfair path, a post-hoc path-specific audit shows that five of eight discovered graphs recover no path from the sensitive attribute to the outcome and therefore report a null overall effect while mediated effects persist, necessitating downstream path-specific fairness analysis along with structural discovery.
发表机构
- University of California, Irvine(加利福尼亚大学欧文分校)
- University of Orléans(奥尔良大学)
机构由 AI 辅助整理,请以论文原文为准。