发表机构
National Chung Hsing University(国立中兴大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模动态动作空间中的一对多移动充电问题,提出LP-BTS学习引导规划架构,结合图提议策略、价值评论器和预算树搜索,在250传感器场景中实现最高存活率0.4545。
AI 中文摘要
许多学习型序贯决策系统直接将当前状态映射到动作。当候选动作数量众多、具有几何结构且随状态重建时,这种捷径会变得脆弱。一对多移动充电使这一场景具体化:在N=250个传感器的情况下,初始状态产生约1,125个候选充电停止动作;每个选定的停止点同时服务于其范围内的传感器,且动作空间随传感器死亡而变化。LP-BTS是一种学习引导的规划架构:图提议策略集中一个小型候选支持集,学习价值评论器评估叶节点,带边预算的PUCT在提交动作前比较短期模拟未来。由于策略在无固定输出头的情况下对该集合评分,单个冻结检查点覆盖所有评估设置,动作空间范围从736到2,813个停止点。匹配的消融实验揭示了互补效应:均匀采样损失8.8个存活率百分点,而在固定目标支持的情况下,PUCT联合保留1.4个百分点(约250个传感器中的3.5个),直接策略选择则多行进23%。在一个前瞻性指定、密封的30场景确认库上评估一次,LP-BTS达到最高观测存活率(0.4545)和存活AUC(0.8031)。其相对于最强领域工程比较器的估计存活优势为+0.0066(95%置信区间[-0.0037, +0.0184]),这是一个未解决的差异,而它在每个配对场景上均超过一个截止时间启发式方法和两个源自源的直接策略重建。两个学习行均为Gong等人报告的变体的源自源重建。在此设置中,结果为大规模动态动作空间中的学习引导规划提供了受控证据。
英文摘要
Many learned sequential decision systems map the current state directly to an action. That shortcut becomes brittle when candidate actions are numerous, geometrically structured, and rebuilt with the state. One-to-many mobile charging makes this setting concrete: with N=250 sensors, the initial state induces about 1,125 candidate charging-stop actions; each chosen stop simultaneously serves its in-range sensors, and the action universe changes as sensors die. LP-BTS is a learning-guided planning architecture: a graph proposal policy concentrates a small candidate support, a learned value critic evaluates leaves, and edge-budgeted PUCT compares short simulated futures before committing an action. Because the policy scores this set without a fixed output head, a single frozen checkpoint covers every evaluated setting, spanning action universes from 736 to 2,813 stops. Matched ablations reveal complementary effects: uniform sampling costs 8.8 survival percentage points, while, with targeted support fixed, PUCT jointly retains 1.4 points (about 3.5 of 250 sensors) and direct policy selection travels 23% farther. On a prospectively specified, sealed 30-scenario confirmatory bank evaluated once, LP-BTS attains the highest observed survival (0.4545) and alive-AUC (0.8031). Its estimated survival advantage over the strongest domain-engineered comparator is +0.0066 (95% CI [-0.0037, +0.0184]), an unresolved difference, while it exceeds a deadline heuristic and two source-derived direct-policy reconstructions on every paired scenario. Both learned rows are trained, source-derived reconstructions of variants reported by Gong et al. In this setting, the results provide controlled evidence about learning-guided planning in a large, dynamic action space.
Comments15 pages, 7 figures. Learning-guided planning, budgeted tree search, PUCT, and sequential decision-making in large dynamic action spaces