arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于样本的先验 elicitation 的最小二乘方法

A Least-Squares Approach to Sample-Based Prior Elicitation

Yannik Pitcan

arXiv 2608.08779首次发表:更新:

发表机构

University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于最小二乘的M估计量,从提供示例点及近似似然的专家处引出贝叶斯先验,建立其渐近性质,放宽假设并通过模拟和真实数据集验证其优于仅样本的最大似然基线。

AI 中文摘要

提供某一量的示例的专家往往会同时表明每个示例的似然程度;我们研究何时该信号值得使用。我们从提供示例点及其近似似然的专家处引出贝叶斯先验,提出通过最小二乘拟合先验——最小化参数密度与引出似然之间的平方偏差,这定义了即使对于不存在矩的分布族也保持适定的M估计量。我们建立了一致性和渐近正态性,并在显式正则条件(包括良好分离和均匀集中要求,我们针对所考虑的分布族验证了这些条件)下,通过将Pinelis关于最大似然估计量的结果扩展到M估计设置,证明了其抽样分布(均匀和非均匀)的非渐近O(1/√n) Berry-Esseen界。随后,我们放宽了实际中限制该方法的最严格假设:专家无需报告密度自身的尺度,未知的报告尺度可通过闭式轮廓化并联合估计,对于位置族无渐近代价。该理论扩展到多元参数,其中方向Berry-Esseen界通过将多元delta方法应用于估计量的平滑隐式代理得到。噪声模型中的加性误差底消除了最优设计的退化,使最优设计处于内部。正态和Beta分布族的模拟证实了预测的n⁻¹误差率和中等样本量下正态近似的准确性。最后,我们将该估计量与仅样本的最大似然基线进行比较,得出专家报告噪声的显式阈值,低于该阈值时引出的似然可证明降低估计误差,并针对11个人类频率判断数据集校准了该阈值。

英文摘要

An expert who supplies examples of a quantity often also signals how plausible each one is; when is that signal worth using? We study eliciting a Bayesian prior from an expert who provides example points together with their approximate likelihoods. We propose fitting the prior by least squares--minimizing the squared discrepancy between a parametric density and the elicited likelihoods--which defines an M-estimator that remains well posed even for families whose moments do not exist. We establish consistency and asymptotic normality, and prove--under explicit regularity conditions, comprising a well-separation and a uniform-concentration requirement that we verify for the families considered--a non-asymptotic O(1/sqrt(n)) Berry--Esseen bound on its sampling distribution, uniform and nonuniform, by extending a result of Pinelis for maximum-likelihood estimators to the M-estimation setting. We then relax the assumptions that most limit the method in practice. Experts need not report on the density's own scale: an unknown reporting scale can be profiled out in closed form and estimated jointly, at no asymptotic cost for location families. The theory extends to multivariate parameters, where a directional Berry--Esseen bound follows from the multivariate delta method applied to a smooth implicit proxy for the estimator. An additive error floor in the noise model removes a degeneracy in the optimal design, making optimal designs interior. Simulations for normal and beta families confirm the predicted n^-1 error rate and the accuracy of the normal approximation at moderate sample sizes. Finally, we compare the estimator with the sample-only maximum-likelihood baseline, derive an explicit threshold on the expert's reporting noise below which the elicited likelihoods provably reduce estimation error, and calibrate that threshold against eleven datasets of human frequency judgments.

Comments34 pages, 7 figures. Code: https://github.com/pitcany/prior-elicitation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑