arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对博弈与学习型对手的信息设计

Information Design Against Gaming and Learning Adversaries

Madhava Gaikwad

arXiv 2609.31643首次发表:更新:

AI 中文总结

针对博弈型与学习型对手,研究二分类器弃权策略的信息设计,揭示两种防御的不可比性,并给出边界重建查询复杂度及帕累托前沿。

AI 中文摘要

一个部署带有弃权(不执行)选项的二分类器的主事者,必须决定机制对哪些查询进行弃权。正确的选择取决于对手的类型。博弈型对手已知分类器,并试图跨越决策边界操纵特征,因此主事者最好的做法是对靠近该边界的查询进行弃权。同样的边界定位规则,对于不知道分类器的学习型对手而言,却是最差的选择:每次弃权现在都告诉对手边界就在附近,这足以驱动一次二分搜索。我们分析这一矛盾。两种自然的防御策略——以固定比率弃权和在边界附近弃权——在Blackwell意义上是不可比较的:任何一种都无法通过后处理另一种的响应来模拟。在第一种防御下,将边界重建到误差$\eps$所需的查询次数为$\tilde\Theta(d/\eps)$,而在第二种防御下为$\Theta(d \log(1/\eps))$,其中$d$是分类器族的VC维,$\tilde\Theta$抑制了关于$d$和$1/\eps$的多对数因子。第一个速率是查询分布上的最坏情况;在达到该速率的分布上,没有任何重建算法能够缩小差距。我们刻画了两种防御目标之间的帕累托前沿,并在七个跨越表格、图像和语言模型特征输入的二元分类任务上确认了这两个速率:标签加反事实访问提取边界所需的查询次数比已发表的仅标签基线最多减少$200\times$。

英文摘要

A principal who deploys a binary classifier with an abstention option must decide which queries the mechanism abstains on. The right choice depends on the adversary. A gaming adversary already knows the classifier and tries to manipulate features across the boundary, so the principal does best by abstaining on queries close to that boundary. The same boundary-localizing rule is the worst possible choice against a learning adversary who does not know the classifier: each abstention now tells the adversary that the boundary is nearby, which is enough to drive a binary search. We analyze this tension. The two natural defenses, abstaining at a fixed rate and abstaining near the boundary, are Blackwell-incomparable: neither can be simulated by post-processing the other's responses. The number of queries needed to reconstruct the boundary to error $\eps$ is $\tildeΘ(d/\eps)$ under the first defense and $Θ(d \log(1/\eps))$ under the second, where $d$ is the VC dimension of the classifier family and $\tildeΘ$ suppresses factors polylogarithmic in $d$ and $1/\eps$. The first rate is a worst case over query distributions; no reconstruction algorithm can close the gap at the distributions that attain it. We characterize the Pareto frontier between the two defense objectives, and confirm both rates on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only baseline.

CommentsAccepted at Gamesec 2026 for Oral

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑