发表机构
Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM交通信号控制仅优化最终结果、无法区分推理步骤优劣的问题,提出ProcessLight框架将决策分解为可验证语义步骤,并基于STeP-PO强化学习实现步骤级信用分配,在真实数据集上验证了优越性。
AI 中文摘要
近年来,大语言模型(LLMs)凭借其在生成人类可读推理方面的优势,被引入作为交通信号控制(TSC)的决策智能体。然而,现有的LLM TSC方法仅从最终结果进行优化,无法区分有效与有缺陷的推理步骤,导致有用或误导性的步骤被共同更新,从而削弱了模型学习有效推理的能力。为弥补这一不足,我们提出了基于LLM的框架ProcessLight,将信号决策分解为可验证的语义步骤。在此基础上,我们进一步开发了逐步交通过程策略优化(STeP-PO),一种新颖的强化学习框架,通过步骤级信用分配来优化结构化推理过程。具体而言,STeP-PO使用步骤质量分数评估局部推理质量,并利用步骤重要性衡量每个步骤对最终动作的影响,然后在语义步骤树结构上分配步骤级优势。由此产生的步骤级优势被传播到推理令牌,实现超越仅结果奖励的细粒度策略优化。在多个真实世界数据集上的广泛实验证明了我们方法的优越性。我们的代码可在以下网址获取:此https URL。
英文摘要
Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readable reasoning generation. Yet, existing LLM TSC methods optimize only from final outcomes and fail to distinguish valid from flawed reasoning steps, causing useful or misleading steps to be jointly updated and thus impairing the model's learning of effective reasoning. To bridge this gap, we propose an LLM-based framework ProcessLight to decompose signal decisions into verifiable semantic steps. Building on ProcessLight, we further develop Step-wise Traffic Process Policy Optimization (STeP-PO), a novel reinforcement learning framework that optimizes structured reasoning processes through step-level credit assignment. Specifically, STeP-PO uses step quality scores to evaluate local reasoning quality and step importance to measure each step's influence on the final action, and then assigns step-level advantages over a semantic step tree structure. The resulting step-level advantages are propagated to reasoning tokens, enabling fine-grained policy optimization beyond outcome-only rewards. Extensive experiments over multiple real-world datasets demonstrate the superiority of our methods. Our code is available at https://github.com/wenzhaoabc/processlight.
CommentsThis paper has been accepted by The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)