arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果逻辑强盗问题中具有反事实公平性的极小极大最优遗憾

Minimax Optimal Regret for Causal Logistic Bandits with Counterfactual Fairness

Junhyuk Huh, Seoungbin Bae, Dabeen Lee

arXiv 2610.01377首次发表:更新:

发表机构

University of Cambridge; KAIST; Seoul National University(剑桥大学; 韩国科学技术院; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对带反事实公平约束的因果逻辑强盗,证明覆盖条件必要性,提出匹配上下界,算法达到极小极大最优遗憾。

AI 中文摘要

我们研究了具有反事实公平性约束的因果逻辑强盗问题。因果结构通过已知的事实和反事实特征映射给出,这些映射共享一个未知的逻辑奖励参数,但学习器仅观察到事实奖励。因此,决定反事实可行性的方向可能无法从可获得的反馈中识别。最接近的先前分析要么省略了覆盖条件,要么施加了相对较强的覆盖条件,并且没有建立匹配的下界。我们首先证明某种覆盖条件是必要的:在没有覆盖类型限制的情况下,具有不同最优公平行动的事实上不可区分的环境会迫使期望联合损失达到$\Omega(T)$。在跨行动池化的事实协方差满足较弱的满秩条件下,我们确定了一个目标特定的信息尺度$V_\star$,它衡量从事实反馈中估计奖励和反事实效果的难度。我们构造了满足此条件的最坏情况族,在该族上每个策略都会产生期望联合损失$\Omega\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}\right)$。我们还给出了一种使用$V_\star$调整的探索-然后-利用过程,以及一种不需要其值的自适应算法。两种算法都实现了$\max\{R_T,V_T\}=\widetilde{O}\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}+\kappa d/\sigma_0^2\right)$,其中$R_T$是相对于最佳公平行动的遗憾,$V_T$表示累积的阶段性正违规。因此,上下界在$T$、$V_\star$和$\min\{\log K,d\}$的主要依赖上匹配,直至对数因子。

英文摘要

We study causal logistic bandits with counterfactual fairness constraints. The causal structure is given through known factual and counterfactual feature maps that share an unknown logistic reward parameter, but the learner observes only factual rewards. Consequently, the directions determining counterfactual feasibility need not be identifiable from the available feedback. The closest prior analyses either omit a coverage condition or impose a comparatively strong one, and do not establish matching lower bounds. We first show that some coverage condition is necessary: without a coverage-type restriction, factually indistinguishable environments with different optimal fair actions force $Ω(T)$ expected joint loss. Under a weaker full-rank condition on the factual covariance pooled across actions, we identify a target-specific information scale $V_\star$ that measures the difficulty of estimating rewards and counterfactual effects from factual feedback. We construct worst-case families satisfying this condition on which every policy incurs expected joint loss $Ω\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}\right)$. We also give an explore--then--exploit procedure tuned using $V_\star$ and an adaptive algorithm that does not require its value. Both algorithms achieve $\max\{R_T,V_T\}=\widetilde{O}\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}+κd/σ_0^2\right)$, where $R_T$ is regret relative to the best fair action and $V_T$ denotes the cumulative stage-wise positive violations. Thus the upper and lower bounds match in their leading dependence on $T$, $V_\star$, and $\min\{\log K,d\}$, up to logarithmic factors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑