结合人类反馈的策略迭代:将后训练强化学习应用于上下文学习
From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation
- Vanderbilt University(范德堡大学)
- Vanderbilt University Medical Center(范德堡大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出结合人类反馈的策略迭代(PIHF)方法,采用预训练语言模型为基底,经实验使其在罕见疾病诊断相关任务中显著提升了多个执行器的Recall@1指标,验证了其可行性。
AI中文摘要:
生成式预训练建立了可复用的任务表征;后续基于语言的任务条件化和上下文学习的研究表明,固定模型可通过指令和演示调整自身行为。结合人类反馈的策略迭代(PIHF)基于这一进展及广义策略迭代的循环评估-改进结构构建,它采用预训练语言模型作为执行基底,将持续修正工作转移至版本化自然语言策略与工具集。语言模型评论家与临床专家审查完整面板推理及工具使用轨迹,以定位循环故障并形成候选修正;专家可重新解释证据,保留采纳与回滚的权限,而Recall@1与Recall@5用于验证候选执行后的结果。在累积消融实验及罕见疾病基准测试中,源自PIHF的策略提升了一个专有执行器及三个参数规模介于30亿至490亿的开放权重执行器的Recall@1,其中GPT-5.4提升32.7个百分点,Qwen3.6-35B提升31.1个百分点,两者差值为1.7个百分点。这些结果支持将预训练语言模型作为固定权重执行基底,用于罕见疾病诊断中专家引导策略开发的可行性。
英文摘要:
Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies. Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively. With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores. In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points). These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models. Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.