arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11135cs.LG

双向多模态融合天空图像与时间序列的大语言模型太阳预报

Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models

  • The University of Melbourne(墨尔本大学)
  • Sri Lanka Institute of Information Technology(斯里兰卡信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge

AI总结:

针对短期光伏预报中云层影响和单模态局限,提出基于大语言模型的双向多模态融合框架SolCloudLLM,融合天空图像与时间序列,在SIRTA和SKIPP'D数据集上显著降低MSE,并在少样本场景中表现最佳。

AI中文摘要:

短期光伏(PV)发电功率和全球水平辐照度(GHI)预报对于有效的调度、备用容量安排和电网运行至关重要。在这些预报时间尺度上,误差主要由云引起的爬坡主导:仅依赖历史数值数据可能难以预测即将到来的云层,这使得地基天空图像成为一种关键的补充物理信号。此外,预报性能对地点和当地观测条件高度敏感,因此对通常稀缺的站点特定数据有强烈需求。近年来,大语言模型(LLMs)在时间序列预测中展现出具有竞争力的性能和较高的数据效率。尽管取得了成功,现有的基于LLM的预测方法仍主要是单模态的,主要依赖历史数值时间序列数据。将天空图像有效纳入基于LLM的预测框架仍是一个探索不足且开放的挑战。在本文中,我们提出了SolCloudLLM,一个基于LLM的多模态预测框架。SolCloudLLM将天空图像块与时间序列块对齐,并通过双向多模态融合融合它们对应的表示,产生一个统一表示,随后映射到LLM的嵌入空间中。在SIRTA和SKIPP'D数据集上的大量实验表明,SolCloudLLM在所有预测时间尺度上始终在MSE指标上优于最佳基线方法,实现了最大相对MSE降低25.4%。分层分析进一步表明,多模态融合的益处主要集中在多云条件下。值得注意的是,SolCloudLLM在几乎所有少样本设置中均取得最佳性能,而其他深度学习基线则经历显著的性能下降,并经常被非学习的物理方法超越。

英文摘要:

Short-term photovoltaic (PV) power and global horizontal irradiance (GHI) forecasts are essential for effective dispatch, reserve scheduling, and grid operations. At these forecasting horizons, errors are predominantly driven by cloud induced ramps: relying solely on historical numerical data may struggle to anticipate an incoming cloud, making ground-based sky images a crucial complementary physical signal. Furthermore, forecast performance is highly sensitive to location and local observing conditions, creating a strong need for site-specific data that are often scarce. Recently, large language models (LLMs) have demonstrated competitive performance and high data efficiency in time-series forecasting. Despite their success, existing LLM-based forecasting methods remain predominantly unimodal, relying primarily on historical numerical time-series data. Effectively incorporating sky imagery into an LLM-based forecasting framework remains under-explored and an open challenge. In this paper, we propose SolCloudLLM, an LLM-based multimodal forecasting framework. SolCloudLLM aligns sky-image patches with time-series patches and fuses their corresponding representations through bidirectional multimodal fusion, yielding a unified representation that is subsequently mapped into the embedding space of an LLM. Extensive experiments on the SIRTA and SKIPP'D datasets demonstrate that SolCloudLLM consistently outperforms the best baseline methods in MSE across all forecasting horizons, achieving a maximum relative MSE reduction of 25.4%. Stratified analysis further indicates that the benefits of multimodal fusion are concentrated primarily under cloudy conditions. Notably, SolCloudLLM achieves the best performance in nearly all few-shot settings, whereas other deep learning baselines experience substantial performance degradation and are frequently outperformed by the non-learning physical method.

↑