arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习合成增强推理的大小-权重边界

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Chengpiao Huang, Kaizheng Wang

arXiv 2608.28576首次发表:更新:

AI 中文总结

该研究针对真实数据稀缺时合成数据引入偏差的问题,提出合成增强推理框架,通过大小-权重边界保证覆盖率,实验中用大语言模型响应增强调查数据,达到目标覆盖率并缩小置信区间。

AI 中文摘要

当真实数据稀缺时,合成数据可改善统计推理,但将合成样本直接当作真实数据处理会引入偏差并导致推理不可靠。我们开发了适用于相关任务总体的合成增强推理通用框架,该框架通过合成观测数量及其权重来刻画合成增强。框架的核心是大小-权重边界,它为每个权重指定了所有更小样本量均能达到目标任务边际覆盖率的最大合成样本量。我们从历史任务中估计该边界,并为估计边界上及下方的所有大小-权重配置同时建立有限样本覆盖率保证。在使用大语言模型响应增强意见调查数据的实验中,我们的方法达到了目标覆盖率,且大幅缩小了置信区间。

英文摘要

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.

Comments19 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑