arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于推断选择信号的贝叶斯方法对超参数敏感性的研究

On the sensitivity to hyperparameters of a Bayesian method to infer selection signatures

Carlos A. Martinez, Ronnal E. Ortiz, Nelson A. Cruz, Carmen H. Cepeda, Juan Sosa

arXiv 2610.03987首次发表:更新:

发表机构

Universidad Nacional de Colombia, Sede Bogotá; Corporación Colombiana de Investigación Agropecuaria, Sede Central; Departament Ciències Matemàtiques i Informàtica de la Universitat de les Illes Balears; Escuela de Matemáticas y Estadística, Universidad Pedagógica y Tecnológica de Colombia, Facultad Seccional Duitama; Departamento de Estadística, Universidad Nacional de Colombia, Sede Bogotá(哥伦比亚国立大学波哥大校区; 哥伦比亚农业研究公司总部; 巴利阿里群岛大学数学与信息科学系; 哥伦比亚教育与技术大学杜伊塔马分校数学与统计学院; 哥伦比亚国立大学波哥大校区统计学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对基于FST的两步贝叶斯推断选择信号方法,通过理论定理与模拟/真实数据实证,揭示了超参数对标记选择及后验估计的敏感性及其渐近行为。

AI 中文摘要

受一个实际问题——推断受选择影响的基因组区域的启发,我们研究了一种基于FST执行此任务的两步方法的若干性质。该方法以其计算简单性作为吸引人的特点。第一步涉及拟合一个共轭贝叶斯模型以获得FST的点估计,并选择在第二步中考虑的标记,在第二步中这些标记使用此类估计进行分组。在此问题中经常使用主观先验,并且可能遇到较大的超参数;因此,解决其影响是有意义的。初步分析表明,当超参数具有相同值且趋于无穷大,或不同但趋于无穷大且其差值趋于零时,第一步中选择的标记数量趋于零。因此,我们正式研究了超参数如何影响后验均值、方差和等位基因频率的预测密度,以及FST的后验分布。此外,还进行了基于模拟和真实数据的实证方法。我们的主要发现总结为两个定理,这些定理设定了方法的渐近行为,为理解上述现象提供了基础。至于实证研究,第二步中创建的组、估计的后验均值和所选标记的数量显示出对超参数的高度敏感性。尽管FST的后验分布在数学上难以处理,但我们的结果揭示了渐近性质,解释了该方法缺乏正式阐明的两个特征。

英文摘要

Motivated by a real-life problem, namely, inferring genomic regions affected by selection, we studied some properties of a two-step method that performs this task based on the FST. This approach exhibits computational simplicity as an appealing feature. Step one involves fitting a conjugate Bayesian model to find point estimates of FST and selecting the markers to be considered in step two, where they are grouped using such estimates. Subjective priors are frequently used in this problem and large hyperparameters may be encountered; hence, there is interest in addressing their impact. Preliminary analyses suggested that, the number of markers selected in step one tends to zero when the hyperparameters have the same value and tend to infinity, or are different but tend to infinity and their difference tends to zero. Thus, we studied how hyperparameters affect the posterior mean, variance, and predictive density of allele frequencies, and the posterior distribution of FST, formally. Moreover, an empirical approach based on simulated and real data was performed. Our main findings are summarized in two theorems that set the asymptotic behavior of the method, providing a foundation to understand the aforementioned phenomena. As to the empirical studies, the groups created in the second step, the estimated posterior means and the number of selected markers showed high sensitivity to hyperparameters. Albeit the mathematical intractability of the posterior distribution of FST, our results revealed asymptotic properties explaining two features of the method that lacked a formal elucidation.

Comments44 pages, 4 figures, 0 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑