AI 中文总结
该研究通过文献计量、统计及NLP多标签文本分析2007-2023年的22695条气候-健康文献记录,明确其增长特征、主题关联规律与方法演变,指出需对缺失注释建模以准确解读时间趋势。
AI 中文摘要
我们分析了2007-2023年的22695条多标签记录构成的经整理的气候-健康文献计量语料库,以表征其增长、主题集中度及演变方法。年度发表量急剧上升,多个变点表明其呈阶段式扩张;负二项模型估计出约10%-11%的同比增长率。暴露-健康共现与独立性存在显著偏离,即使在考虑边缘术语流行度后,典型的危害-结局二元组(如极端高温与热相关影响、洪水/飓风与心理健康)出现频率仍远高于预期。针对哮喘标签记录的分层逻辑模型显示其与空气污染相关暴露(包括臭氧和颗粒物)高度一致,而通用热/温度术语的占比相对不足。方法层面,时间尺度建模随时间向更长跨度转变,同时至少一种传统方法标签有所下降。最后,我们检测到时间和地理依赖的注释完整性,包括近年暴露术语编码减少,这凸显了在解读时间趋势时需对缺失值进行建模的必要性。
英文摘要
We analyzed a curated climate--health bibliographic corpus of 22,695 multi-labeled records from 2007--2023 to characterize growth, thematic concentration, and evolving methods. Annual publication counts rose sharply, with multiple change-points indicating phase-structured expansion; Negative Binomial models estimated roughly 10--11% year-over-year growth. Exposure--health co-occurrence departed strongly from independence, with canonical hazard--outcome dyads (e.g., extreme heat with heat-related impacts; floods/hurricanes with mental health) occurring far more often than expected even after accounting for marginal term popularity. A hierarchical logistic model for asthma-tagged records showed strong alignment with air-pollution-related exposures (including ozone and particulate matter) and relative under-representation of generic heat/temperature terms. Methodologically, modeling timescales shifted toward longer horizons over time, while at least one legacy method tag declined. Finally, we detected time- and geography-dependent annotation completeness, including decreased exposure-term coding in recent years, underscoring the need to model missingness when interpreting temporal trends.
Comments29 Pages, 13 figures