On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
在LLM规划中的泛化差距:测试与验证者奖励强化学习
机构 * Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genoa(信息学、生物工程、机器人学和系统工程系,热那亚大学) ; AIKO S.r.l.(AIKO公司)
专题命中 规划推理 :planning(title,abstract);verifier(title,abstract);分类 cs.AI、cs.LG
AI总结 本研究发现微调LLM在规划任务中存在显著的泛化差距,通过三种诊断干预揭示模型依赖领域特定模式而非可转移能力。
Comments 9 pages, 4 figures, 3 tables, 2 pages of supplementary materials. Submitted to a conference implementing a double-blind review process