RLMEval: Evaluating Research-Level Neural Theorem Proving
专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025 Findings. RLMEval benchmark released: https://github.com/augustepoiroux/RLMEval
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025 Findings. RLMEval benchmark released: https://github.com/augustepoiroux/RLMEval