Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL
向内看,向外探:通过分层强化学习从LLM内部状态学习温度策略
机构 * Zhejiang University(浙江大学) ; Southeast University(东南大学) ; The Chinese University of Hong Kong(香港中文大学) ; Shanghai Innovation Institute(上海创新研究院)
AI总结 Introspective LLM通过分层强化学习从LLM内部状态学习温度策略,提升数学推理任务中的探索效率与可解释性。