arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

部分观测传播森林中病原体感染群体的人口统计学推断

Demographic inference of pathogen-infected populations from partially observed transmission forests

Matthew Hall

arXiv 2609.23624首次发表:更新:

发表机构

London School of Hygiene and Tropical Medicine(伦敦卫生与热带医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用部分观测传播森林中孤立个体的信息,通过闭式概率模型推断病原体感染群体的人口统计学参数,并应用于功效计算、采样假设检验及人口推断。

AI 中文摘要

基因组流行病学已有多种方法用于识别哪些采样个体在传播链中通过近距离相关联,无论是通过遗传距离阈值将其分组为簇,还是通过识别可能的直接传播对。被发现没有采样邻居的个体——即孤立个体——通常被搁置一旁。我们认为它们具有信息价值。任意两个采样个体是否被证明相关联,取决于采样框架的大小以及独立谱系引入其中的数量,因此关联与未关联个体的平衡本身即是关于这些数量的数据。我们通过将传播树与采样框架的交集视为一个带标签的有根森林来形式化这一点,从该森林中均匀随机采样固定数量的节点。利用全子式矩阵树定理,我们推导出在N个节点上,指定节点集合为独立集的带标签有根k-森林数量的闭式表达式。由此,我们再次以闭式形式获得给定大小的样本完全不包含关联个体的概率,以及任意观测到的簇配置的似然性,既包括每个簇内传播结构已知的情形,也包括仅知道簇大小的情形。该方法更接近于调查或标记-重捕法,而非传统的系统动力学模型拟合:信息仅来自单一样本的关联结构,且不对病原体动态做任何假设。我们概述了三个应用:前瞻性研究的功效计算、对将簇解读为传播热点所依据的均匀采样假设的检验,以及人口统计学推断本身。最后,我们列出了更灵活的实现所需放宽的假设。

英文摘要

Genomic epidemiology has several methods for identifying which sampled individuals are linked by close proximity in the transmission chain, whether by grouping them into clusters under a genetic distance threshold or by identifying probable direct transmission pairs. Individuals found to have no sampled neighbours---singletons---are usually set aside. We argue that they are informative. Whether any two sampled individuals prove to be linked depends on the size of the sampling frame and the number of independent lineage introductions into it, and thus the balance of linked and unlinked individuals is itself data about these quantities. We formalise this by treating the intersection of a transmission tree with a sampling frame as a labelled rooted forest, from which a fixed number of nodes are sampled uniformly at random. Using the all-minors matrix-tree theorem, we derive a closed-form expression for the number of labelled rooted $k$-forests on $N$ nodes in which a specified set of nodes is independent. From this we obtain, again in closed form, the probability that a sample of a given size contains no linked individuals at all, and the likelihood of an arbitrary observed configuration of clusters, both when the transmission structure within each is known, and when only the cluster sizes are. The approach is closer to a survey, or to mark-recapture, than to conventional phylodynamic model fitting: the information comes from the linkage structure of a single sample alone with no assumptions regarding pathogen dynamics. We outline three applications: power calculations for prospective studies, a test of the uniform sampling assumption that underlies the reading of clusters as transmission hotspots, and demographic inference itself. We finally set out the assumptions that a more flexible implementation would need to relax.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑