Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
支付更少的泛化税:RL训练对LLM代理跨域泛化能力的研究
机构 * Meta Superintelligence Labs(Meta超智能实验室) ; FAIR at Meta(Meta的FAIR) ; Northwestern University(西北大学) ; The Ohio State University(俄亥俄州立大学) ; University of Pennsylvania(宾夕法尼亚大学)
AI总结 本研究探讨了RL训练对LLM代理跨域泛化能力的影响,发现增加状态信息丰富度可提升泛化性能,同时指出建模选择对泛化能力的关键作用。