当暴露-结局关系为非线性时,双样本孟德尔随机化估计的是什么?
What does two-sample Mendelian randomization estimate when the exposure-outcome relationship is nonlinear?
- Department of Statistics and Data Science, The Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院统计与数据科学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文推导了双样本孟德尔随机化在暴露-结局关系非线性时的估计含义,指出其比率估计为剂量-反应曲线斜率的加权平均,且合并估计量缺乏固定解释,异质性检验和MR-Egger会误判非线性为多效性,强调线性假设需明确辩护并优先使用个体水平数据。
AI中文摘要:
基于汇总统计数据的双样本孟德尔随机化已成为流行病学中一种流行的孟德尔随机化设计。其比率估计仅在包括线性在内的强不可检验假设下具有平均因果效应解释。我们推导了当连续暴露具有非线性效应时,每个变异的比率以及由此构建的合并估计量和稳健估计量所估计的内容,这基于两个关于遗传变异和未测量因素如何塑造暴露的条件:不存在未测量的共同效应修饰因子,以及存在暴露控制函数。在任一条件下,每个变异的比率是剂量-反应曲线斜率的加权平均值,其权重取决于该变异如何移动暴露分布,并且当不同基因型的暴露分布不交叉时,这些权重为非负。使每个人的暴露均等移动的变异所针对的量几乎相同;作用于特定剂量或暴露离散程度的变异则针对不同的量,其中一些超出了任何个体效应的范围。因此,合并估计量没有固定的解释,异质性检验和MR-Egger将非线性误读为多效性,而汇总统计数据无法揭示该问题,即使剂量-反应曲线可从个体水平数据中识别。对于连续暴露,双样本估计依赖于应被明确说明和辩护的线性假设;个体水平暴露数据能解决该问题,应予以优先考虑。
英文摘要:
Two-sample Mendelian randomization with summary statistics has become a popular MR design in Epidemiology. Its ratio estimate has an average causal effect interpretation only under strong untestable assumptions including linearity. We derive what the per-variant ratio, and the pooled and robust estimators built from it, estimate when a continuous exposure has a nonlinear effect, under two conditions on how genetic variation and unmeasured factors shape exposure: no unmeasured common effect modifier, and an exposure control function. Under either condition, each variant's ratio is a weighted average of the slope of the dose-response curve, with weights that depend on how that variant moves the distribution of exposure and that are nonnegative when the exposure distributions at different genotypes do not cross. Variants that shift everyone's exposure equally target nearly the same quantity; variants acting at particular doses or on the spread of exposure target different ones, some outside the range of any individual effect. Pooled estimates therefore have no fixed interpretation, heterogeneity tests and MR-Egger misread nonlinearity as pleiotropy, and summary statistics cannot reveal the problem, even when the dose-response curve is identifiable from individual-level data. For continuous exposures, two-sample estimates rest on a linearity assumption that should be stated and defended; individual-level exposure data resolve the problem and should be prioritized.