arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

行星预测引擎:通过智能数据选择和基础模型嵌入实现自主地理空间预测

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun, Hamed Sadeghi, Dav M. Ebengo, Mbulayi Onesime, Ciara Judge, Rouslan Solomakhin, John Wamburu, William Ogallo, Aisha Walcott-Bryant, Sanxing Chen, Arbaaz Muslim, Yael Mayer, Ronald Ho, Roy Lee, Ruth Alcantara, Abdoulaye Diack, Monica Bharel, Lambert Rosique, Jeremy Amez-Droz, Christopher Haire, James Manyika, Yossi Matias, Niv Efron, Gautam Prasad, Shravya Shetty

arXiv 2608.26088首次发表:更新:

发表机构

Google Research; Institut National de Recherche Biomédicale(谷歌研究院; 国家生物医学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出行星预测引擎(PPE),一种自主AI系统,可根据自然语言查询完成地理空间预测的端到端工作流程,在多类任务中优于现有模型,降低了行星规模分析的技术门槛。

AI 中文摘要

应对粮食安全、灾害风险、疾病暴发和社会经济脆弱性等关键全球挑战,需要高保真地理空间建模。然而,构建预测性行星模型仍受碎片化数据生态系统的制约,该过程需要手动数据检索、多模态数据整理与融合,以及迭代式模型选择。我们提出行星预测引擎(Planetary Prediction Engine, PPE),这是一种自主AI系统,可直接根据自然语言查询执行上述端到端工作流程。PPE动态合成多模态数据集,从开放网络和地球观测平台(Data Commons、Google Earth Engine)检索时空相关协变量,并将其与地理空间基础模型嵌入(PDFM、AlphaEarth)融合;同时,它通过自动过拟合防护机制,搜索针对任务定制的模型架构族。在不同任务、地理区域和科学领域中,PPE始终优于最先进模型或人工调优的专家基线。对于美国空间回归任务,PPE在21项CDC健康指标上的平均R²为76.8%(基线为60.0%),在FEMA国家风险指数上为64.9%(基线为60.0%),在社会脆弱性指数上为66.2%(基线为58.6%)。在数据稀缺环境下的空间降尺度任务中,PPE整合局部代理,使尼日利亚粮食安全指标的基线准确率翻倍(R²为66.1%,基线为31.5%)。对于2026年刚果民主共和国本巴地区埃博拉疫情的流行病学即时预测,PPE的Recall@10为83.3%(在五次每周预测中识别出18个新入侵卫生区中的15个),较公开最先进模型提升了10.3个百分点(约73%)。通过将自主多模态行星数据发现与针对性模型优化相结合,PPE降低了行星规模分析的技术门槛,支持快速、定制化、专家级别的部署。

英文摘要

Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI system that executes this end-to-end workflow directly from natural-language queries. PPE synthesizes multimodal datasets on the fly, retrieving spatiotemporally relevant covariates across open-web and Earth observation platforms (Data Commons, Google Earth Engine) and fusing them with geospatial foundation model embeddings (PDFM, AlphaEarth). Simultaneously, it searches over task-tailored model architecture families with automated overfitting guards. Across diverse tasks, geographies, and scientific domains, PPE consistently outperforms state-of-the-art or manually tuned expert baselines. For US spatial regression, PPE improves mean $R^2$ across 21 CDC health indicators (76.8% vs. 60.0%), FEMA national risk indices (64.9% vs. 60.0%), and the Social Vulnerability Index (66.2% vs. 58.6%). For spatial downscaling in data-scarce settings, PPE integrates localized proxies to double baseline accuracy in Nigerian food security indicators ($R^2$ of 66.1% vs. 31.5%). For epidemiological nowcasting of the 2026 DRC Bundibugyo Ebola outbreak, PPE achieves a Recall@10 of 83.3% (identifying 15 of 18 newly invaded health zones across five weekly forecasts), a +10.3 percentage-point improvement over the public state-of-the-art modeling (~73%). By combining autonomous multimodal planetary data discovery with targeted model optimization, PPE lowers the technical barrier to planetary-scale analytics, enabling rapid, customized, expert-level deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑