Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL
向内看,向外探:通过分层强化学习从LLM内部状态学习温度策略
机构 * Zhejiang University(浙江大学) ; Southeast University(东南大学) ; The Chinese University of Hong Kong(香港中文大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 Introspective LLM通过分层强化学习从LLM内部状态学习温度策略,提升数学推理任务中的探索效率与可解释性。