arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

NVIDIA(英伟达)

2026-02-05 至 2026-02-05 共收录 3
2501.06148 2026-02-05 cs.LG stat.ML

From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training

从离散时间策略到连续时间扩散采样器:渐近等价与更快的训练

Julius Berner, Lorenz Richter, Marcin Sendera, Jarrid Rector-Brooks, Nikolay Malkin

机构 * California Institute of Technology(加州理工学院) NVIDIA(英伟达) Zuse Institute Berlin(柏林泽尼克研究所) dida Datenschmiede GmbH(dida数据隐私公司) Jagiellonian University(雅盖隆大学) Mila, Université de Montréal(蒙特利尔大学机器学习研究所) University of Edinburgh(爱丁堡大学) CIFAR Fellow, Learning in Machines and Brains(CIFAR Fellow, 机器学习与大脑学习)

AI总结 本文提出通过渐近等价性将离散时间策略转化为连续时间扩散采样器,提升训练效率和采样性能。

Comments TMLR final version; code: https://github.com/GFNOrg/gfn-diffusion/tree/stagger

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04056 2026-02-05 eess.SY cs.RO cs.SY

Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World

模块化安全护栏是现实世界中由基础模型赋能的机器人所必需的

Joonkyung Kim, Wenxi Chen, Davood Soleymanzadeh, Yi Ding, Xiangbo Gao, Zhengzhong Tu, Ruqi Zhang, Fan Fei, Sushant Veer, Yiwei Lyu, Minghui Zheng, Yan Gu

机构 * Purdue University(普渡大学) Amazon(亚马逊公司) NVIDIA(英伟达公司)

AI总结 本文提出模块化安全护栏以解决现实世界中由基础模型赋能的机器人在开放、长尾和随时间适应的环境中的安全挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14843 2026-02-05 cs.LG cs.AI cs.CL

The Invisible Leash: Why RLVR May or May Not Escape Its Origin

无形的绳索:为何RLVR可能或可能不逃脱其起源

Fang Wu, Weihao Xuan, Ximing Lu, Mingjie Liu, Yi Dong, Zaid Harchaoui, Yejin Choi

机构 * stanford(斯坦福大学) tokyo(东京大学) nvidia

AI总结 RLVR可能限制模型发现原创解决方案,其在提升精度的同时可能缩小探索范围,需未来创新以扩展代表性不足的解决方案区域。

详情

展开后加载摘要…

URL PDF HTML 收藏