arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

F-DACE:面向弃权安全型对话式零售决策支持的模糊分歧感知因果证据融合

F-DACE: Fuzzy Disagreement-Aware Causal Evidence Fusion for Abstention-Safe Conversational Retail Decision Support

Sourish Dey

arXiv 2609.18238首次发表:更新:

发表机构

Centric Software(Centric Software)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

F-DACE提出一种模糊分歧感知的因果证据融合决策层,通过硬性否决实现弃权安全,在模拟和零售应用中降低错误推荐率,并支持对话式决策支持。

AI 中文摘要

观测性决策支持系统即使在合理的估计器之间存在分歧时,也常常仅展示一个因果估计作为推荐。所提出系统的内在引擎是因果机器学习:通过后门调整识别条件平均处理效应估计量,由EconML DML因果森林和DoWhy线性回归进行估计,通过双向固定效应进行检验,并通过约束优化转化为候选杠杆。F-DACE是该引擎上的决策层。它将精度、倾向性重叠、安慰剂反驳稳定性、区间重叠和方向一致性表示为模糊隶属度。硬性否决在估计量不匹配、诊断失败、信息性符号冲突或证据薄弱时强制弃权(不执行)。在跨越六种识别条件的180个面板模拟中,F-DACE在67.2%的运行中做出决策,并将错误推荐限制在17.2%;相应的比率对于因果森林为33.3%,对于后门回归为35.6%,与确定性一致投票相匹配而非超越它。几乎所有(31个中的30个)错误推荐都发生在共享未测量混杂下,当每个组件都共享遗漏变量时,任何融合规则都无法诊断。零售应用将公开的Walmart面板聚合为45家商店的6,435个商店-周。F-DACE对所有五个降价指标均弃权(不执行):一些估计不精确,一个反驳失败,且MarkDown5存在直接符号冲突。一个LangGraph对话代理展示影响、假设和杠杆优化工具,而确定性验证器保持因果层状态。在24个实时问题上,它实现了100.0%的工具路由准确率、100.0%的状态保真度和0.983的平均扎根性。在十个对抗性问题上,它抵抗了所有注入的指令。

英文摘要

Observational decision-support systems often expose one causal estimate as a recommendation even when plausible estimators disagree. The inherent engine of the proposed system is causal machine learning: a conditional-average-treatment-effect estimand identified by backdoor adjustment, estimated by an EconML DML causal forest and DoWhy linear regression, checked by two-way fixed effects, and converted into candidate levers by constrained optimisation. F-DACE is the decision layer on that engine. It represents precision, propensity overlap, placebo-refutation stability, interval overlap, and directional agreement as fuzzy memberships. Hard vetoes force abstention after estimand mismatch, failed diagnostics, informative sign conflict, or weak evidence. In 180 panel simulations spanning six identification conditions, F-DACE made a decision in 67.2% of runs and limited false recommendations to 17.2%; the corresponding rates were 33.3% for the causal forest and 35.6% for backdoor regression, matching deterministic unanimity rather than dominating it. Nearly all (30 of 31) false recommendations occurred under shared unmeasured confounding, which no fusion rule can diagnose when every component shares the omitted variable. The retail application aggregates a public Walmart panel to 6,435 store-weeks across 45 stores. F-DACE abstains for all five markdown indicators: some estimates are imprecise, one refutation fails, and MarkDown5 has a direct sign conflict. A LangGraph conversational agent exposes impact, what-if, and lever-optimization tools while a deterministic verifier preserves causal-layer status. On 24 live questions it achieved 100.0% tool-routing accuracy, 100.0% status fidelity, and 0.983 mean groundedness. On ten adversarial questions it resisted all injected instructions.

CommentsPages: 21,Figures: 6,Tables: 10

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑