从原子到智能体:迈向大语言模型智能体数学能力的可解释评估
From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities
另 1 家 · 查看机构详情
- Sun Yat-sen University(中山大学)
- Tencent Youtu Lab(腾讯优图实验室)
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出过程级基准,对齐智能体行为与数学原子能力分类,评估LLMs智能体数学推理能力,发现相似端到端准确率的模型智能体能力轮廓差异显著,凸显过程级评估的重要性。
中文摘要 AI 辅助
大语言模型(LLMs)正从执行端到端数学推理向整合智能体智能演进,但现有多数数学基准仅评估最终答案。这种结果导向的评估对识别过程级故障或严谨逻辑的诊断价值有限,无法指导LLMs向鲁棒智能体转型。为填补这一空白,本文提出一种过程级基准,用于评估LLMs内在的智能体数学推理能力。我们的框架将解题智能体行为与可复用数学原子能力的结构化分类对齐,设计了涵盖文本和多模态场景的规划、行动、反馈任务综合套件,该套件由自动化流水线支撑,可生成高质量轨迹并通过受控LLM重写产生细粒度注释。实验显示,具有相似端到端准确率的模型可呈现显著不同的智能体能力轮廓,这表明过程级评估对解释LLMs的真实潜力及指导下一代数学智能体开发至关重要。
英文摘要
Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagnostic value for identifying process-level failures or rigorous logic, failing to guide the transformation of LLMs into robust agents. To bridge this gap, we present a process-level benchmark designed to evaluate the inherent agentic mathematical reasoning abilities of LLMs. Our framework aligns problem-solving agentic behaviors with a structured taxonomy of reusable mathematical atomic capabilities. We design a comprehensive suite of planning, action, and feedback tasks across both textual and multimodal contexts, supported by an automated pipeline that synthesizes high-quality trajectories and produces fine-grained annotations via controlled LLM rewriting. Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles. This demonstrates that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.