发表机构
NVIDIA; Fudan University(英伟达; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出大型离散策略(LDiP),通过随机迭代评分从物理合理候选动作中显式选择,在自动驾驶和机器人操作中优于或匹敌连续生成策略,提供可解释的行为建模方案。
AI 中文摘要
行为策略通常被表述为连续生成模型,其迭代去噪过程具有表现力,但难以解释且容易产生不合理的动作。我们提出大型离散策略(LDiP),一种完全离散的行为建模框架,从大量物理上合理的候选动作中选择动作。LDiP并非扰动动作,而是通过随机迭代评分来提高表现力:它逐步重新评分并利用评分空间中的随机性对候选动作进行剪枝,从而在保留显式决策过程的同时,实现对合理动作的细粒度排序和探索。在端到端规划、闭环驾驶、机器人操作和视觉-语言-动作设置中,LDiP在自动驾驶方面持续优于强大的离散和连续基线,并在机器人操作方面达到或超过连续生成策略。这些结果表明,当配备有效的评分机制时,离散策略为行为建模提供了一种有表现力、合理且可解释的替代方案。项目网站:此HTTPS URL。
英文摘要
Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vocabulary of physically plausible candidates. Rather than perturbing actions, LDiP improves expressivity through stochastic iterative scoring: it progressively re-scores and prunes candidates with score-space stochasticity, enabling fine-grained ranking and exploration among plausible actions while preserving an explicit decision process. Across end-to-end planning, closed-loop driving, robotic manipulation, and vision-language-action settings, LDiP consistently outperforms strong discrete and continuous baselines in autonomous driving, and exceeds or matches continuous generative policies in robotic manipulation. These results show that discrete policies, when equipped with effective scoring mechanisms, offer an expressive, plausible, and interpretable alternative for behavior modeling. Project website: https://zhenxinli.net/LargeDiscretePolicy/.