仅精确动作值不够:针对多区域VAV控制的推理模型的展开验证强化微调
Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
查看机构详情
- The University of Tokyo(东京大学)
- Tokyo University of Science(东京理科大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对多区域VAV控制,测试前沿LLM的控制能力及TD3引导的RFT效果,发现RFT未产生持续改进,需先开展状态转移聚焦的监督微调。
中文摘要 AI 辅助
多区域变风量(VAV)控制必须在多个连续执行器间平衡热舒适度、室内空气质量和电力使用。模型预测控制与强化学习被广泛研究,但部署通常需要针对特定建筑的建模或训练,限制了可扩展性。我们首先测试前沿推理模型(一种经训练可利用额外推理时计算的大语言模型,即LLM)能否无需特定建筑训练,仅从文本实现有竞争力的VAV控制;确认该能力后,我们测试TD3引导的强化微调(RFT)能否将控制知识迁移至可本地部署的开放权重模型。在基于物理的四区域模拟器中,对三个夏季日评估了五种控制器:与基于Guideline 36的基线相比,TD3降低了4.5%的HVAC电力消耗,同时提升了温度与CO₂合规性;无需特定建筑训练的GPT-5实现了最大降幅(6.2%),但降低了通风余量。对于RFT,确定性展开会恢复保存的状态、应用一个候选方案并遵循TD3对每个动作评分;审计学习到的评判器与这些展开的对比,发现了其近完美跨时间相关性(r=0.9998)所掩盖的缺陷:状态内排名不可靠,该评判器仅在10个状态中的5个选出了展开最优的候选。即便有展开验证,200步RFT也未在采样动作回报上产生持续改进;开放权重控制器训练前后的电力消耗均高于基线,其5分钟预测仍逊于 persistence(持续性)。GPT-5对状态转移的预测远更准确;精确展开评分虽能对采样动作排名,但既未揭示下一个状态的影响也未指明改进方向。未变的状态转移误差促使在基于价值的RFT前,开展以状态转移为重点的监督微调。
英文摘要
Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model predictive control and reinforcement learning are widely studied, but deployment typically requires building-specific modeling or training, limiting scalability. We first test whether a frontier reasoning model (an LLM trained to use additional inference-time computation) can achieve competitive VAV control from text without building-specific training. With that capability established, we then test whether TD3-guided reinforcement fine-tuning (RFT) can transfer control knowledge into a locally deployable open-weight model. Five controllers are evaluated over three summer days in a physics-based four-zone emulator. Relative to a Guideline 36-based baseline, TD3 reduced HVAC electricity by 4.5% while improving temperature and CO$_2$ compliance. Without building-specific training, GPT-5 achieved the largest reduction (6.2%) but reduced the ventilation margin. For RFT, deterministic rollouts restore a saved state, apply one candidate, and follow TD3 to score each action. Auditing a learned critic against these rollouts exposed a failure hidden by its near-perfect across-time correlation ($r=0.9998$): within-state ranking was unreliable; the critic selected the rollout-best candidate in only 5 of 10 states. Even with the rollout verifier, 200 RFT steps produced no sustained improvement in sampled-action return; the open-weight controller used more electricity than the baseline before and after training, and its five-minute predictions remained worse than persistence. GPT-5 predicted transitions far better. Exact rollout scores rank sampled actions but reveal neither next-state effects nor an improvement direction. The unchanged transition errors motivate transition-focused supervised fine-tuning before value-based RFT.