FLARE:基于小样本学习的自适应反思引擎
FLARE: Few-shot Learning-based Adaptive Reflective Engine
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出FLARE框架,利用小样本学习与反思机制优化指令,在多基准任务中优于GEPA,数据效率更高且稳定性更强,凸显小样本学习策略优化对提升LLM性能的重要性。
AI中文摘要:
大型语言模型(LLM)正越来越多地被部署在复杂的复合AI系统中,其性能取决于提示词的质量。近期的最先进优化器如GEPA(遗传-帕累托算法)提出,反思式指令演化的性能优于传统强化学习和小样本优化。本研究通过引入FLARE(Few-shot Learning-based Adaptive Reflective Engine,基于小样本学习的自适应反思引擎)框架,对这一观点提出挑战,该框架利用先进的反思机制和少量小样本参考示例来优化指令。我们在多样化的基准套件上评估了该方法,涵盖检索增强推理(HotPotQA、MedQA、2WikiMultiHopQA)、工具调用以及多标签情感分类(GoEmotions),使用GPT-5系列模型。结果表明,FLARE在所有任务-模型对中均持续优于GEPA:在HotPotQA上,使用GPT-5-Chat时FLARE比GEPA高出14.2个百分点(52.2 vs 42.2);工具调用任务中FLARE达到87.0%,而GEPA为81.0%;在GoEmotions的完整5408样本测试集上,使用GPT-5.1时FLARE将微F1提升至52.7%,较GEPA的提升幅度(+5.7)高出两倍以上。除原始准确率外,FLARE的数据效率也显著更高:在GoEmotions任务中,FLARE仅使用100个验证样本即可达到峰值性能,且在随机种子下的稳定性明显优于GEPA。研究结论表明,尽管反思式指令功能强大,但小样本学习的策略性优化仍是最大化下一代LLM潜力的关键前沿领域。
英文摘要:
Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-of-the-art optimizers like GEPA (Genetic-Pareto) have argued that reflective instruction evolution can outperform traditional reinforcement learning and few-shot optimization. In this work, we challenge this shift by introducing FLARE (Few-shot Learning-based Adaptive Reflective Engine), a framework that leverages advanced reflective mechanisms and a small set of few-shot reference examples to optimize instructions. We evaluate our method across a diverse suite of benchmarks -- spanning retrieval-augmented reasoning (HotPotQA, MedQA, 2WikiMultiHopQA), tool calling, and multi-label emotion classification (GoEmotions) -- using the GPT-5 series of models. Our results demonstrate that FLARE consistently outperforms GEPA, winning on every task-model pair: it achieves gains of up to +14.2 points on HotPotQA (52.2 vs. GEPA's 42.2 with GPT-5-Chat), reaches 87.0% on tool calling (vs. 81.0% for GEPA), and lifts GoEmotions micro-F1 to 52.7% (+15.3) with GPT-5.1 on the full 5408-example test split, more than doubling GEPA's +5.7 gain. Beyond raw accuracy, FLARE is also strikingly data-efficient: on GoEmotions it reaches its peak performance using as few as 100 validation examples, while remaining markedly more stable across random seeds than GEPA. Our findings suggest that while reflective instructions are powerful, the strategic optimization of few-shot learning remains a critical frontier for maximizing the potential of next-generation LLMs.