arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于观察性因果推断的人类增强智能工作流程

A Human-Augmenting Agentic Workflow for Observational Causal Inference

Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Nathan Kallus

arXiv 2607.22443首次发表:更新:

AI 中文总结

研究针对观察性因果推断任务,介绍开源Python包`oci-agent`,其实现人在回路智能工作流程,自动化繁琐环节,支持多种处理效应估计,经案例研究和评估性能良好,在Netflix广泛应用。

AI 中文摘要

数据分析智能体在应用和科研中愈发常见。但对于观察性因果推断这类高度专业化任务,仍需人工监督确保结果有效性。我们引入了开源Python包`oci-agent`,它实现了用于观察性因果推断的人在回路智能工作流程。它旨在自动化应用因果推断中重要但繁琐的方面,如协变量平衡检查等,让人类能专注更细致任务。2026年6月首次开源时支持单个二元处理平均处理效应的双稳健学习,之后又增加了对异质处理效应估计和通过部分线性模型进行多个连续处理的支持。文中描述了`oci-agent`背后的原理,并给出了Netflix内部案例研究及对其能力的公开数据评估。在众多评估中,`oci-agent`优于结构较少的基线,与手工调整的基准保持竞争力。它在Netflix被广泛用于因果推断,自发布以来每月编排超100次分析。

英文摘要

Data analysis agents are becoming increasingly common tools for applied and scientific research. Yet, for highly specialized tasks such as Observational Causal Inference (OCI), human oversight remains necessary to ensure the validity of results. We introduce `oci-agent`, an open-source Python package that implements a human-in-the-loop agentic workflow for observational causal inference. `oci-agent` is designed to automate vital but laborious aspects of applied causal inference, such as covariate balance checking, propensity score trimming, and sensitivity analysis, so that humans can focus on more nuanced tasks, such as framing questions, scrutinizing assumptions, and evaluating diagnostics and results. We initially open-sourced `oci-agent` in June 2026 with support for doubly robust learning of the average treatment effect of a single binary treatment. Since then, we have added support for heterogeneous treatment effect estimation and for multiple continuous treatments via partially linear models. In this paper, we describe the principles behind `oci-agent` and offer internal Netflix case studies and evaluations on public data of its capabilities. Across numerous evaluations, `oci-agent` outperforms less structured baselines while remaining competitive with hand-tuned benchmarks. `oci-agent` is used extensively for causal inference at Netflix and has orchestrated more than 100 analyses per month since its release in June.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑