arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12856cs.LG

基于验证器的热能存储控制推理模型强化微调

Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control

  • The University of Tokyo(东京大学)
  • Tokyo University of Science(东京理科大学)

机构由 AI 辅助整理,请以论文原文为准。

Takumi Shioda, Kohei Terashima, Tatsuo Nagai

AI总结:

研究针对建筑物热能存储控制中模型难以扩展的问题,采用基于可验证奖励的强化学习,通过强化微调训练开放权重推理模型,在简单办公楼TES基准测试中取得较好效果,明确了规划模式转移情况,推动了相关测试与验证器发展。

AI中文摘要:

建筑物需要根据电网状况转移冷却负荷,热能存储(TES)可实现此转移,但需在存储约束下提前数小时规划。模型预测控制(MPC)和强化学习难以在多建筑物中扩展。本研究通过具有可验证奖励的强化学习(RLVR)来调整开放权重推理模型。将精确的离线动态规划(DP)动作值转换为每个候选动作的密集奖励。仅用30个训练提示,强化微调(RFT)将模型训练为上层调度器,从基于文本的状态和预测中输出每小时热泵设定点。评估使用一个故意简化的办公楼TES基准,其中精确DP易于处理且最优解已知。RFT将开放权重模型的排放量从70.5降至61.2千克二氧化碳,接近DP最优值60.8千克二氧化碳。GPT-5在无特定任务训练的情况下几乎与DP和MPC匹配,而非推理的语言模型GPT-4o产生的排放量高于无存储基线,因此推理时的推理似乎很重要。跟踪分析表明RFT主要稳定可观察的规划模式(候选比较、前瞻和可行性检查)而非创建新策略。稳健性和泛化测试明确了哪些可转移:强化的规划模式在预测误差和未见的TES条件下持续存在,并可转移到电池任务,但不同结构限制了收益。基于DP的可验证奖励为使开放权重推理模型适应建筑物存储调度提供了实用方法。这些结果推动了对全建筑控制的更高保真测试以及城市规模能源管理的可扩展验证器的发展。

英文摘要:

Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well requires planning hours ahead under storage constraints. Model predictive control (MPC) and reinforcement learning are difficult to scale across buildings. This study instead adapts an open-weight reasoning model through reinforcement learning with verifiable rewards (RLVR). We convert exact offline dynamic-programming (DP) action values into dense rewards for every candidate action. Using only 30 training prompts, reinforcement fine-tuning (RFT) trains the model as an upper-level scheduler that outputs hourly heat-pump setpoints from text-based states and forecasts. Evaluation uses a deliberately simple office-building TES benchmark where exact DP is tractable and the optimum is known. RFT reduces the open-weight model's emissions from 70.5 to 61.2 kg-CO2, close to the DP optimum of 60.8 kg-CO2. GPT-5 nearly matches DP and MPC without task-specific training, while GPT-4o, a non-reasoning LLM, produces higher emissions than the no-storage baseline, so inference-time reasoning appears important. Trace analysis shows that RFT mainly stabilizes observable planning patterns (candidate comparison, look-ahead, and feasibility checking) rather than creating a new strategy. Robustness and generalization tests clarify what transfers: the reinforced planning patterns persist under forecast errors and an unseen TES condition and carry over to a battery task, but its different structure limits the gains. DP-based verifiable rewards offer a practical way to adapt open-weight reasoning models to building storage scheduling. These results motivate higher-fidelity tests of whole-building control and scalable verifiers for city-scale energy management.

补充信息

↑