Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
神经概念验证器:通过概念编码扩展证明者-验证者博弈
Berkant Turan, Suhrab Asadulla, David Steinmann, Kristian Kersting, Wolfgang Stammer, Sebastian Pokutta
机构
*
Zuse Institute Berlin, Germany
;
AI \& ML Lab, Computer Science Department, TU Darmstadt
;
Hessian Center for AI (hessian.AI)
;
Lab1141
;
German Research Center for AI (DFKI)
;
Institute of Mathematics, Technische Universit\"at Berlin, Germany
;
Max Planck Institute for Informatics, SIC
Comments28 pages, 5 figures, 12 tables, revised references. An earlier version of this work was presented at the ICML 2025 Workshop on Actionable Interpretability
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable)
DPO 解绑:你的训练算法在人类选择理论中实际上是解耦的(且其损失函数的凸性并非必需)
Wenxuan Zhou, Shujian Zhang, Brice Magdalou, John Lambert, Ehsan Amid, Richard Nock, Andrew Hard
机构
*
Google Research(谷歌研究)
;
Google DeepMind(谷歌DeepMind)
;
Work done while at Google DeepMind(在谷歌深Mind的工作)
;
CEE-M, Univ Montpellier, CNRS, INRAE, Institut Agro(CEE-M,蒙彼利埃大学,CNRS,INRAE,农业研究所)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
Chain-of-Goals 分层策略用于长视界离线目标条件强化学习
Jinwoo Choi, Sang-Hyun Lee, Seung-Woo Seo
机构
*
Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea(首尔国立大学电气与计算机工程系)
;
Department of Automotive Engineering, Ajou University, Gyeonggi-do, South Korea(全州大学汽车工程系)
机构
*
Department of Statistics and Data Science(统计与数据科学系)
;
Department of Statistics(统计系)
;
Data Science, Southern University of Science(数据科学,南方科技大学)
;
College of Computing(计算学院)
;
College of Computing and Data Science(计算与数据科学学院)
;
School of Statistics(统计学系)
;
School of Statistics and Data Science(统计与数据科学系)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(人工智能学院,香港中文大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
Learning Bug Context for PyTorch-to-JAX Translation with LLMs
学习 PyTorch 到 JAX 翻译中的 Bug 上下文
Hung Phan, Son Vu, Tuan Dinh, Nesreen Ahmed, Ali Payani, Ali Jannesari
机构
*
Department of Computer Science, Iowa State University, Ames, IA, USA(Iowa State大学计算机科学系)
;
Department of Human Science, Iowa State University, Ames, IA, USA(Iowa State大学人类科学系)
;
Cisco AI Research, Outshift, San Jose, CA, USA(Cisco AI研究,Outshift,San Jose,加利福尼亚州,美国)
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis
无模型鲁棒平均奖励强化学习及样本复杂度分析
Zachary Roch, George Atia, Yue Wang
机构
*
Department of Electrical and Computer Engineering, University of Central Florida, Orlando, Florida, USA(电气与计算机工程系,中央佛罗里达大学,奥兰多,佛罗里达州,美国)
;
Department of Computer Science, University of Central Florida, Orlando, Florida, USA(计算机科学系,中央佛罗里达大学,奥兰多,佛罗里达州,美国)
Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
通过增强潜在本征属性将重光照作为视觉先验的探针
Xiaoyan Xing, Xiao Zhang, Sezer Karaoglu, Theo Gevers, Anand Bhattad
机构
*
UvA-Bosch Delta Lab, University of Amsterdam, Amsterdam, Netherlands(乌得勒支大学阿姆斯特丹分校博世Delta实验室)
;
The University of Chicago, Chicago, USA(芝加哥大学)
;
Johns Hopkins University, Baltimore, USA(约翰霍普金斯大学)
MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
MEAL: 持续多智能体强化学习基准
Tristan Tomilin, Luka van den Boogaard, Samuel Garcin, Constantin Ruhdorfer, Bram Grooten, Fabrice Kusters, Yali Du, Andreas Bulling, Mykola Pechenizkiy, Meng Fang
机构
*
Eindhoven University of Technology, The Netherlands(埃因霍温理工大学,荷兰)
;
University of Edinburgh, UK(爱丁堡大学,英国)
;
University of Stuttgart, Germany(斯图加特大学,德国)
;
King's College London, UK(伦敦国王学院,英国)
;
University of Liverpool, UK(利物浦大学,英国)