arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20268cs.AIcs.CL

PoTRE:受认知异质性启发的测试时推理

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

Anmol Kankariya, Sercan Ö. Arık

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型复杂推理难题,提出PoTRE异构框架,将推理解耦为四个智能体,经任务自适应聚合层协调,在三个前沿基准测试中评估,在HLE上达最优准确率,以相似或更少令牌提升推理性能。

中文摘要 AI 辅助

虽然大语言模型在许多任务中表现出色,但在需要长期规划和迭代纠错的复杂推理上常常遇到困难。此外,当模型遇到新抽象或严格领域约束时,标准单流提示很脆弱。我们引入了PoTRE(多拓扑推理集成),这是一个异构框架,将推理解耦为四个智能体:对抗细化智能体、分层战略规划智能体、频谱搜索智能体和直接链智能体。一个最终的任务自适应聚合层通过最终候选选择、语义合成或神经符号验证动态协调这些视角,以产生一个强大的全局解决方案。我们在三个前沿基准上评估了PoTRE:ARC-AGI-2、人类的最后一场考试(HLE)和PRBench Finance。PoTRE在HLE上达到了49.92%的当前最优准确率,超过了之前的最佳官方分数。我们证明,与大量扩展的同构基线相比,这种架构异质性在使用相似或更少推理令牌的情况下实现了更高的推理性能。

英文摘要

While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives -- via final candidate selection, semantic synthesis, or neuro-symbolic verification -- to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines.

补充信息

↑