arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06331cs.ROcs.SYeess.SY

几何分布控制:利用部分结构知识的学习进展

Geometric Distributional Control: Learning Progress with Partial Structural Knowledge

Tong Wu

首次发表
浏览论文内容

中文总结 AI 辅助

针对部分结构知识下的实时控制,提出几何分布控制(GDC),将控制分解为可行性与进展,利用进展加权数据学习局部价值梯度,在多层优化和驾驶任务中优于现有方法。

中文摘要 AI 辅助

实时控制通常处于两种限制性机制之间。当动力学、参数、目标和在线规划模型被指定时,预测优化和基于模型的控制是强大的;强化学习可以放宽这一要求,但必须从序列数据和交互中推断长期价值信号,这使得训练缓慢、方差高,并且在大型动作空间中难以扩展。这种中间机制在包括自动驾驶、仓储机器人、交通控制和配送无人机在内的系统中很常见:部分几何、物理、规则或约束是已知的,但任务进展的局部方向仍然不确定。几何分布控制(GDC)专为这种部分知识设置而设计。它将控制分解为可行性和进展:已知的几何、规则、约束和响应映射定义了一个可执行的脚手架,而进展加权的可行数据在该脚手架上学习缺失的方向信号。学习到的分数充当类似贝尔曼的局部价值梯度,选择能够取得进展的动作,而无需全局贝尔曼递归、完全指定的规划器或同时吸收可行性和偏好的黑盒策略。这种知识可以是轻量级和部分的,例如简单的动力学、安全过滤器、局部地图、约束投影器或较低级别的响应映射;它不需要编码完整的动力学或长期目标。离线时,GDC从短的已知可行片段中拟合一个进展倾斜的分布,并带有弱的带符号进展证书。在线时,其分数通过脚手架投影并以滚动时域反馈的方式应用。我们证明了该分数会下降一个数据诱导的软进展价值,并在结构化多层优化和SUMO路线进展驾驶上验证了GDC,在这些场景中,它优于仅已知求解器和学习基线,同时保持脚手架强制执行的可行性。

英文摘要

Real-time control often sits between two limiting regimes. Predictive optimization and model-based control are powerful when dynamics, parameters, objectives, and online planning models are specified; reinforcement learning can relax this requirement, but must infer long-horizon value signals from sequential data and interaction, making training slow, high-variance, and hard to scale in large action spaces. This middle regime is common in systems including autonomous driving, warehouse robotics, traffic control, and delivery drones: partial geometry, physics, rules, or constraints are known, yet the local direction of task progress remains uncertain. Geometric Distributional Control (GDC) is designed for this partial-knowledge setting. It factorizes control into feasibility and progress: known geometry, rules, constraints, and response maps define an executable scaffold, while progress-weighted feasible data learns the missing directional signal on that scaffold. The learned score acts as a Bellman-like local value-gradient, selecting actions that make progress without requiring global Bellman recursion, a fully specified planner, or a black-box policy that absorbs both feasibility and preference. This knowledge can be lightweight and partial, such as simple dynamics, safety filters, local maps, constraint projectors, or lower-level response maps; it need not encode full dynamics or a long-horizon objective. Offline, GDC fits a progress-tilted distribution from short known-feasible snippets with weak signed progress certificates. Online, its score is projected through the scaffold and applied in receding-horizon feedback. We prove that this score descends a data-induced soft progress value and validate GDC on structured multilevel optimization and SUMO route-progress driving, where it improves over known-only solvers and learning baselines while preserving scaffold-enforced feasibility.

发表机构

  • University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

↑