AI 中文总结
该研究将新闻级联视为文章数量与可信度构成的耦合随机过程,以拉斯维加斯枪击案为案例拟合对比多类模型,发现混合计数模型与马尔可夫类可信度模型拟合效果最优,且方法论可推广至其他事件但结果随事件特征变化。
AI 中文摘要
我们将由单个高影响力事件触发的新闻级联视为一对耦合随机过程进行研究:每日文章数量$N(t)$以及可信与不可信内容的构成$\pi_{\mathrm{pos}}(t)$。我们以2017年拉斯维加斯枪击案为主要案例研究(采用CC-News数据集与NewsGuard工具,涵盖91天、2702篇文章),在统一的似然尺度下,拟合并比较了针对$N(t)$的四种计数过程模型,以及针对$\pi_{\mathrm{pos}}(t)$的四种分布模型。在数量建模方面,非齐次泊松(IHP)加霍克斯(Hawkes)的混合模型在AIC、BIC和对数似然指标上均取得压倒性优势,其中IHP用于解释初始的外生冲击,而Hawkes模型用于解释后续的自激发现象。对Hawkes衰减率进行的剖面似然扫描显示,最优Hawkes核可简化为AR(1)单步滞后形式$\lambda(t)=\mu+nN(t-1)$。数量层面的这种马尔可夫结构直接启发了第二部分的可信度模型:我们将可信度标签视为二元状态空间上的一阶马尔可夫链,并使用共享的条件二项似然,将齐次、区制转换和平均场参数化模型与双变量Hawkes基线模型进行对比。结果表明,双变量Hawkes模型和非齐次马尔可夫模型的拟合效果几乎无法区分,因为它们已穷尽信号信息,开始拟合噪声。针对2017年哈维飓风数据的简要复现实验表明,同一方法论适用,但由于哈维飓风是逐渐内生成因的事件,因此得出了不同的结果。
英文摘要
We study the news cascade triggered by a single high-impact event as a pair of coupled stochastic processes: the daily article count $N(t)$ and the reliable-vs-unreliable composition $π_{\mathrm{pos}}(t)$. Using the 2017 Las Vegas shooting as a primary case study (CC-News + NewsGuard, 91 days, 2,702 articles), we fit and compare four counting-process models for $N(t)$ and four distributional models for $π_{\mathrm{pos}}(t)$ on a common likelihood scale. For counts, a hybrid Inhomogeneous-Poisson plus Hawkes model wins decisively on AIC, BIC, and log-likelihood, as the IHP accounts for the initial exogenous shock, while the Hawkes accounts for the subsequent self-excitation. A profile-likelihood sweep of the Hawkes decay rate collapses the optimal Hawkes kernel to an AR(1) one-step lag $λ(t)=μ+nN(t-1)$. This Markovian structure on the counts directly motivates the Part-B reliability models: we treat the reliability label as a first-order Markov chain on a binary state space and test homogeneous, regime-switching, and mean-field parameterizations against a bivariate Hawkes baseline model using a shared conditional binomial likelihood. The results show that the bivariate Hawkes model and the non-homogeneous Markov models are nearly indistinguishable on fit, as they exhaust the signal and start fitting to the noise. A brief replication on the 2017 Hurricane Harvey data shows that the same methodology applies but produces different results due to Hurricane Harvey's gradual endogenous build-up.
Comments16 pages, 8 figures, 7 tables