能源视觉-语言-行动:面向意图条件住宅能源管理的受控多模态基准
Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management
AI总结:
本文提出EVLA基准,将住宅电池调度构建为多模态轨迹预测问题,通过受控实验揭示语言路径对预测质量的关键作用,而视觉路径无显著误差优势。
AI中文摘要:
视觉-语言-行动(VLA)模型主要在机器人领域得到研究,其中视觉观察和语言指令被映射为物理动作。本文介绍了能源视觉-语言-行动(EVLA),一个面向意图条件住宅能源管理的受控多模态基准。EVLA将电池调度构建为多模态轨迹预测问题,其中RGB能量场表示、数值运行状态和自然语言目标被映射为由有限时域基于采样的参考生成器生成的16步电池动作轨迹。源窗口来源于公共住宅电力负荷数据,而电价、电池荷电状态、室内温度和一天中的时间作为基准元数据生成。一个隐藏的运行机制仅通过能量场纹理编码,使得在显式数值状态固定的情况下能够实现成对的视觉变化。将439,203个保留的基础窗口与三种隐藏机制和五种语言目标交叉,产生6,588,045个多模态实例。一项初步研究使用固定的5,000个训练、500个验证和500个测试实例子集,在三个训练种子上评估了36种配置。在MobileNet系列比较中,移除处理后的语言使轨迹均方误差从0.3856 +/- 0.0039增加到0.8628 +/- 0.0001,而移除视觉则得到0.3843 +/- 0.0013,与完整模型相当。结果表明模态使用存在强烈不对称性:处理后的语言路径与预测质量密切相关,而当前RGB路径不提供总体误差优势。这些结果表征了固定的试点子集和执行的协议,而非完整基准训练。EVLA为研究语义意图和潜在上下文如何影响住宅能源动作预测提供了一个受控环境。
英文摘要:
Vision-Language-Action (VLA) models are studied mainly in robotics, where visual observations and language instructions are mapped to physical actions. This paper introduces Energy Vision-Language-Action (EVLA), a controlled multimodal benchmark for intent-conditioned residential energy management. EVLA frames battery scheduling as a multimodal trajectory-prediction problem in which an RGB energy-field representation, a numerical operating state, and a natural-language objective are mapped to a 16-step battery-action trajectory generated by a finite-horizon sampling-based reference generator. Source windows are derived from public residential electrical-load data, while electricity price, battery state of charge, indoor temperature, and time of day are generated benchmark metadata. A hidden operating regime is encoded only through energy-field texture, enabling paired visual changes while the explicit numerical state is fixed. Crossing 439,203 retained base windows with three hidden regimes and five language objectives yields 6,588,045 multimodal instances. An initial study evaluates 36 configurations over three training seeds using fixed subsets of 5,000 training, 500 validation, and 500 test instances. In the MobileNet-family comparison, removing processed language increases trajectory mean-squared error from 0.3856 +/- 0.0039 to 0.8628 +/- 0.0001, whereas removing vision yields 0.3843 +/- 0.0013, comparable to the full model. The results show strong asymmetry in modality use: the processed-language pathway is strongly associated with prediction quality, while the current RGB pathway provides no aggregate error advantage. These results characterize the fixed pilot subset and executed protocol rather than full-benchmark training. EVLA provides a controlled setting for studying how semantic intent and latent context influence residential energy-action prediction.