MIRAGE:基于强化学习引导的多视角创意语言模型推理
MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MIRAGE提出一种推理时多视角创意推理框架,通过选择器与推理器结合,在多个基准上以低开销提升LLM复杂任务准确率。
AI中文摘要:
近年来,大型语言模型(LLMs)的进步彻底改变了人工智能以及人类与AI的交互方式。尽管取得了令人瞩目的进展,LLMs在处理复杂的数学、科学和逻辑任务时仍面临困难。受人类认知灵活性——即我们动态切换心理视角的能力——的启发,我们提出了MIRAGE(通过智能体引导探索的多视角推理时推理),一种新颖的推理时创意思维框架。MIRAGE包含一个选择器(Selector),用于优先考虑有效的概念视角(例如代数、概率),以及一个推理器(Reasoner),它依次解决任务,直到出现自信的解决方案,否则聚合多个视角。在GSM8K、MATH500、MMLU-Pro和Game-of-24基准测试上,MIRAGE始终优于诸如思维链和多样化提示集成等方法,以最小的推理开销显著提高了准确性,为实际应用提供了可扩展的解决方案。
英文摘要:
Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility - our ability to dynamically switch mental perspectives - we propose MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration), a novel inference-time creative thinking framework. MIRAGE includes a Selector that prioritizes effective conceptual perspectives (e.g., algebraic, probabilistic) and a Reasoner that sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives. Tested on GSM8K, MATH500, MMLU-Pro, and Game-of-24 benchmarks, MIRAGE consistently outperforms methods like Chain-of-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications.