arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProxyGuard:针对带共享目标的随机数据发布机制的直接可靠性推断

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

Dipesh Tharu Mahato, Pramod Dhungana

arXiv 2608.18643首次发表:更新:

发表机构

New York University; Queens College(纽约大学; 皇后学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ProxyGuard通过两种模式控制随机数据发布机制的错误,在注册研究中提升了直接模式的功效,且可审计Rice--TVAE等机制。

AI 中文摘要

研究人员常从众多发布结果、变换或种子中选择代理数据集。搜索可能使无效发布显得充分,而一个充分的发布并不能证明其生成器可靠。ProxyGuard利用预先指定的有界风险和密封目标集控制两类错误。命名发布模式校正多重性并认证特定发布;直接共享目标模式评估公共目标上的独立机制抽样,对其有利评分率求下界,并减去无效发布贡献的有利评分的界。在给定目标的条件下,发布评分相互独立,无需独立目标批次或发布层面p值依赖假设,即可得到有限样本机制可靠性保证。我们证明仅均值惩罚是尖锐的,并推导了带有加性目标浓度的平滑评分证书。在一项注册三要求研究中,直接模式在可靠性0.95时将功效从5.6%提升至64.2%,而命名模式在高信号证据下仍更优。前瞻性审计覆盖了全流程Rice--TVAE(每次抽样都重新训练)和非表格文本机制。

英文摘要

Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard controls both errors using prespecified bounded risks and a sealed target set. Named-release mode corrects for multiplicity and certifies specific releases. Direct shared-target mode evaluates independent mechanism draws on a common target, lower-bounds their favorable-score rate, and subtracts a bound on favorable scores contributed by invalid releases. Conditional on the target, release scores are independent, yielding a finite-sample mechanism-reliability guarantee without independent target batches or assumptions on release-level $p$-value dependence. We show that the mean-only penalty is sharp and derive a smooth-score certificate with additive target concentration. In a registered three-requirement study, direct mode raises power from 5.6\% to 64.2\% at reliability 0.95, while named mode remains stronger under high-signal evidence. Prospective audits span full-pipeline Rice--TVAE, which retrains on every draw, and a non-tabular text mechanism.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑