arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11662cs.AI

拉马克驾驶学校:通过进化竞争发现自动驾驶训练策略

Lamarck's Driving School: Discovering Autonomous Driving Training Strategies through Evolutionary Competition

Yichun Ye, He Zhang, Ye Tian, Jian Sun

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出拉马克进化框架,通过候选分布竞争发现自动驾驶训练策略,完整框架可降性能损失最多25.07%,轻量策略仍降19.13%,还揭示了训练分布价值变化的阶段性规律。

中文摘要 AI 辅助

自动驾驶能力高度依赖训练过程中遇到的场景分布。现有方法通常使用真实度、难度或风险等替代标准来构建或动态调整训练场景分布,但这些预定义的替代标准可能会错误表征训练价值,导致训练资源利用效率低下。为解决这一局限,我们提出一种拉马克进化框架,用候选分布间的竞争替代基于替代标准的引导。我们将训练策略发现建模为多阶段双层优化问题,并使用拉马克进化算法近似求解:外层通过达尔文式交叉、变异和选择探索场景分布空间;内层通过策略学习获取新能力,再通过拉马克遗传将其传递至后续阶段,使场景分布与策略能力协同进化。所得进化轨迹揭示了高价值分布中反复出现的阶段性规律,其特征是通过阶段性挑战轮换实现能力积累。我们进一步将这些规律提炼为轻量、可复用的拉马克训练策略。实验表明,与基线相比,完整框架可将性能损失降低多达25.07%,而轻量策略仍能实现19.13%的降低。这些结果证明,进化竞争既能发现有效训练策略,又能揭示训练分布价值随策略能力变化的可复用阶段性模式。代码已在GitHub上开源。

英文摘要

Autonomous driving capabilities depend strongly on the distribution of scenarios encountered during training. Existing methods commonly construct or dynamically adapt training scenario distributions using surrogate criteria such as realism, difficulty, or risk. However, these predefined surrogates may misrepresent training value, leading to inefficient use of training resources. To address this limitation, we propose a Lamarckian evolutionary framework that replaces surrogate-based guidance with competition among candidate distributions. We formulate training strategy discovery as a multi-stage bilevel optimization problem and use Lamarckian evolution algorithm to approximate its solution. At the outer level, Darwinian crossover, mutation, and selection explore the scenario distribution space; at the inner level, policy learning acquires new capabilities, and Lamarckian inheritance transfers them to subsequent stages, allowing scenario distributions and policy capabilities to co-evolve. The resulting evolutionary trajectories reveal recurring stage-wise regularities among high-value distributions, characterized by capability accumulation through stage-wise challenge rotation. We further distill these regularities into a lightweight, reusable Lamarckian Training Strategy. Experiments show that, compared with the baseline, the complete framework reduces performance loss by up to 25.07%, while the lightweight strategy still achieves a 19.13% reduction. These results demonstrate that evolutionary competition can both discover effective training strategies and reveal reusable stage-wise patterns in how the value of training distributions changes with policy capability. Code is available on GitHub.

↑