arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

NVIDIA(英伟达)

2026-08-24 至 2026-08-24 共收录 2
2508.14313 2026-08-24 cs.LG cs.AI 版本更新

AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning

你的强化学习奖励函数是你的最佳搜索PRM:统一强化学习与基于搜索的文本生成

Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang, Zihan Dong, Marco Pavone, Ligong Han, Zhang-Wei Hong, Tong Che, Dimitris N. Metaxas

机构 * Rutgers University(新泽西州立大学) Nanyang Technological University(南洋理工大学) University of Connecticut(康涅狄格大学) Fudan University(复旦大学) NVIDIA Research(NVIDIA研究) Red Hat AI Innovation(红帽AI创新) MIT-IBM Watson AI Lab(MIT-IBM沃森AI实验室) Massachusetts Institute of Technology(麻省理工学院)

AI总结 本文提出AIRL-S,通过强化学习与搜索技术的统一,利用奖励函数直接学习动态PRM,提升推理链扩展和跨任务泛化能力,实验显示性能提升9%并优于基线PRM。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21624 2026-08-24 cs.LG physics.chem-ph physics.comp-ph 版本更新

HIP: Hessian Interatomic Potentials without derivatives

HIP从臀部射出:无需导数的Hessian互作用势

Andreas Burger, Luca Thiede, Nikolaj Rønne, Varinia Bernales, Nandita Vijaykumar, Tejs Vegge, Arghya Bhowmik, Alan Aspuru-Guzik

机构 * University of Toronto(多伦多大学) Vector Institute for Artificial Intelligence(向量人工智能研究所) Technical University of Denmark(丹麦技术大学) CAPeX Pioneer Center for Accelerating P2X Materials Discovery(CAPeX加速P2X材料发现先锋中心) Acceleration Consortium(加速联盟) Canadian Institute for Advanced Research (CIFAR)(加拿大高等研究院) NVIDIA(英伟达)

AI总结 本文提出通过深度学习模型直接预测Hessian,无需自动微分或有限差分,实现更高效、准确的分子力学计算。

Comments this https URL (https://github.com/BurgerAndreas/hip)

详情

展开后加载摘要…

URL PDF HTML 收藏