arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10507 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10507 篇

1911.00262 2019-11-04 cs.LG cs.CL cs.IR stat.ML 81%

Finding the most similar textual documents using Case-Based Reasoning

Marko Mihajlovic, Ning Xiong

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.04076 2019-09-13 cs.CL cs.AI 81%

Counterfactual Story Reasoning and Generation

Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, Yejin Choi

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.05656 2019-08-16 cs.LG cs.AI stat.ML 81%

PHYRE: A New Benchmark for Physical Reasoning

Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, Ross Girshick

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.10550 2018-11-27 cs.AI cs.CL 81%

Challenges in the Automatic Analysis of Students' Diagnostic Reasoning

Claudia Schulz, Christian M. Meyer, Michael Sailer, Jan Kiesewetter, Elisabeth Bauer, Frank Fischer, Martin R. Fischer, Iryna Gurevych

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.05457 2018-03-16 cs.AI cs.CL cs.IR 81%

Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 7 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01063 2026-08-25 cs.AI 版本更新 80%

MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

MindClaw: 用于精确干预的闭环具身心理状态推理

Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Wen-Huang Cheng, Jianlong Fu

机构 * Jilin University(吉林大学) Microsoft Asia(微软亚洲) National Taiwan University(国立台湾大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 提出MindClaw框架,通过闭环具身心理状态推理实现精确干预,结合多源输入、信念记忆、认知触发技能和动作生成,在动态环境中优化干预时机。

Comments Extended version of the CVPR 2026 paper *MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents*. This work is in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13872 2026-08-24 cs.NE cs.AI 版本更新 80%

S-AI-Recursive: Convergent Recursive Reasoning

S-AI-Recursive:一种生物启发式且时间稀疏的AI架构,用于迭代、反思和节能推理

Said Slaoui

机构 * Mohammed V University(穆莱·伊斯梅尔大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出S-AI-Recursive架构,通过激素闭环迭代而非单向传递实现推理,结合生物启发式方法和数学模型,验证了时间稀疏性原理。

Comments Preprint. 55 pages. No figures. S-AI-Recursive: Convergent Recursive Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09743 2026-07-14 cs.AI cs.GT 新提交 80%

Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets

为策略制定者搭建框架:霍特林空间市场中与架构相关的推理干预

Pratyush Singh

机构 * California Institute of Technology(加州理工学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 研究结构化推理干预对大语言模型战略经济推理的影响及与模型架构的关系,以霍特林模型评估GPT - 4.1 - mini和GPT - 5 - mini,发现支架类型与模型架构有交叉交互作用,对抗性测试有损害,还存在陈述性 - 程序性差距。

Comments 26 pages (11 main + 15 appendix), 6 figures, 4 tables. Accepted at the ICLR 2026 Workshop on LLM Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29254 2026-06-30 cs.CL cs.NE 80%

Travel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge Graphs

面向旅游的推理大语言模型:基于领域特定知识图谱

Vignesh Ram Nithin Kappagantula, Shayan Hassantabar, Samuel Simpson, Golnaz Moallem

机构 * Expedia Group(Expedia集团)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 提出一种模块化流水线,利用专家构建的知识图谱生成多跳问答对,微调大语言模型以提升旅游领域推理的准确性和校准能力,在基准上达到82.4%精确匹配。

Comments Accepted to the Uncertainty Reasoning and Quantification in Decision Making (UDM) Workshop, KDD 2026 (To be presented in August 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10479 2026-06-10 cs.AI 新提交 80%

ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics

ComBench: 奥林匹克级组合数学中严格证明推理与构造实现的基准测试

Shunkai Zhang, Haoran Zhang, Yun Luo, Qianjia Cheng, Haodi Lei, Yizhuo Li, Runzhe Zhan, Zhilin Wang, Bangjie Xu, Yucheng Su, Xinmiao Han, Xiaoye Qu, Dongrui Liu, Zhouchen Lin, Yu Qiao, Ning Ding, Yafu Li, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) Peking University(北京大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 提出ComBench基准,包含100道奥林匹克级组合问题,分分析和构造两类,通过评分与验证评估大模型推理能力,发现最强模型准确率仅65.4%,且证明推理与构造实现能力存在差异。

Comments 39 pages, 6 figures, 26 tables. Project page: https://simplified-reasoning.github.io/ComBench/docs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07966 2026-04-22 cs.CV cs.CL 80%

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

视觉-表格问答:面向表格图像推理的开放领域基准

Boammani Aser Lompo, Marc Haraoui

机构 * École de Technologie Supérieure(埃克塞技术学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文提出Visual-TableQA,一个大规模开放领域多模态数据集,用于评估和提升视觉推理能力,包含2500个LaTeX渲染表格和6000对推理密集型问答对,通过多模型协作生成,验证了模型在外部基准上的鲁棒性。

Comments Accepted at the First Workshop on Foundations of Reasoning in Language Models, NeurIPS 2025. Available at: https://openreview.net/forum?id=fvJRsGwhPf

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08140 2026-04-10 cs.CR cs.AI cs.MM cs.NI 80%

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

基于LLM的多模态推理用于加密流量解释:一个基准

Longgang Zhang, Xiaowei Fu, Fuxiang Huang, Lei Zhang

机构 * School of Microelectronics and Communication Engineering, Chongqing University(重庆大学微电子与通信工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出BGTD基准和mmTraffic框架,通过结合原始字节与结构化注释,实现可解释的加密流量解释,生成高保真的人可读报告,同时保持高分类准确率。

Comments Project page \url{https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark}

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12896 2026-03-30 cs.CL 80%

None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks

并非其他:一种区分推理与记忆的通用技术,用于多选LLM评估基准

Eva Sánchez Salido, Julio Gonzalo, Guillermo Marco

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文提出一种通用方法,通过改变数学问题的数值来区分LLM的推理能力与记忆能力,评估了多个模型在公开和私有数据集上的表现,发现模型在该方法下准确率显著下降,揭示了记忆在当前LLM回答中的重要作用。

Journal ref "On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks," in IEEE Access, vol. 14, pp. 9384-9393, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19669 2026-03-30 cs.AI 80%

HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization

HeaRT:一种基于分层电路推理树的代理框架用于AMS设计优化

Souradip Poddar, Chia-Tung Ho, Ziming Wei, Weidong Cao, Haoxing Ren, David Z. Pan

机构 * ECE Department, The University of Texas at Austin(德克萨斯大学奥斯汀分校电子与计算机工程系) NVIDIA Corporation(英伟达公司) The George Washington University(乔治华盛顿大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 HeaRT提出了一种分层电路推理树的代理框架,通过提升F1(subcircuits)和F1(loops)指标,实现更高效的AMS设计优化,且在不同架构上表现出更好的适应性和收敛速度。

Comments Analog Design Automation, Hierarchical Circuit Reasoning, Context-Aware Design Adaptation, LLMs, Agentic Frameworks, Electronic Design Automation (EDA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22389 2026-02-18 cs.DL cs.AI 80%

Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?

小模型和推理大模型能否对期刊文章进行科研质量评分?平均和少样本学习是否有帮助?

Mike Thelwall, Ehsan Mohammadi

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文评估了小模型和推理模型对期刊文章科研质量评分的能力,发现4b以上的小模型在使用评分平均时表现良好,但推理模型无明显优势。

Comments Thelwall, M. & Mohammadi, E. (2026). Can small and reasoning Large Language Models score journal articles for research quality and do averaging and few-shot help? Scientometrics

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13401 2026-01-21 cs.CV cs.AI 80%

Reasoning with Pixel-level Precision: QVLM Architecture and SQuID Dataset for Quantitative Geospatial Analytics

基于像素级精度的推理:QVLM架构与SQuID数据集用于定量遥感分析

Peter A. Massih, Eric Cosatto

机构 * Department of Machine Learning, NEC Laboratories America(机器学习系,NEC美国实验室)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出QVLM架构和SQuID数据集,通过解耦语言理解和视觉分析,提升定量空间推理的准确性。

Comments Submitted to CVPR 2026. Introduces the QVLM architecture and the SQuID dataset for quantitative geospatial reasoning. Dataset DOI: 10.57967/hf/7565

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02794 2025-11-05 cs.AI cs.MA 80%

When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning

Chenyu Zhang, Minsol Kim, Shohreh Ghorbani, Jingyao Wu, Rosalind Picard, Patricia Maes, Paul Pu Liang

机构 * Harvard University(哈佛大学) MIT Media Lab(麻省理工学院媒体实验室)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments Accepted at the Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10207 2025-10-15 cs.AI 80%

Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning

Yujian Zhang, Keyu Chen, Zhifeng Shen, Ruizhi Qiao, Xing Sun

机构 * Tencent Youtu Lab(腾讯优图实验室)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02892 2025-10-06 cs.LG 80%

RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning

Aleksei Arzhantsev, Otmane Sakhi, Flavian Vasile

机构 * Criteo AI Lab(Criteo人工智能实验室) Ecole Polytechnique Paris(巴黎高等理工学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

Comments Accepted to the Efficient Reasoning Workshop at NeuRIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19676 2025-09-18 cs.AI 80%

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

Lachlan McGinness, Peter Baumgartner

机构 * School of Computer Science, Australian National University and CSIRO(计算机科学学院,澳大利亚国立大学和CSIRO)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments The original version of this article was withdrawn because there were errors in the evaluation of model faithfulness to reasoning strategies and completeness of reasoning. The analysis was re-conducted correctly and version two contains the corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10541 2025-07-16 cs.CL 80%

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Zhuoshi Pan, Qizhi Pei, Yu Li, Qiyao Sun, Zinan Tang, H. Vicky Zhao, Conghui He, Lijun Wu

机构 * Tsinghua University(清华大学) OpenDataLab, Shanghai Artificial Intelligence Laboratory(开放数据实验室、上海人工智能实验室) Renmin University of China(中国人民大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments REST (Reasoning Evaluation through Simultaneous Testing), a stress-testing framework that concurrently exposes LRMs to multiple problems simultaneously

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12001 2023-10-25 cs.CL 80%

OPT-R: Exploring the Role of Explanations in Finetuning and Prompting for Reasoning Skills of Large Language Models

Badr AlKhamissi, Siddharth Verma, Ping Yu, Zhijing Jin, Asli Celikyilmaz, Mona Diab

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Proceedings of the 1st Workshop on Natural Language Reasoning and Structured Explanations (NLRSE) at ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.01205 2020-07-21 cs.CL 80%

CS-NLP team at SemEval-2020 Task 4: Evaluation of State-of-the-art NLP Deep Learning Architectures on Commonsense Reasoning Task

Sirwe Saeedi, Aliakbar Panahi, Seyran Saeedi, Alvis C Fong

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments 6 pages, 1 figure, 2 tables, SemEval -2020, Commonsense Reasoning and Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05822 2026-08-07 cs.SE 新提交 80%

Agent-Based Test Assertion Generation via Diverse Perspective Aggregation

基于智能体的多视角聚合测试断言生成

Dong Wang, Qiaoyu Han, Lin Yang, Jianyi Zhou, Guangtai Liang, Junjie Chen

专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 针对现有LLM断言生成方法的局限,提出基于智能体的AssertMate框架,通过三个组件聚合多视角,在Defects4J和EvoSuite验证中性能显著优于现有技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04788 2026-06-26 cs.SE 版本更新 80%

Adaptive Intellect Unleashed: The Feasibility of Knowledge Transfer in Large Language Models

自适应智能释放:大型语言模型中知识迁移的可行性

Qing Huang, Yishun Wu, Zhenchang Xing, He Jiang, Yu Cheng, Huan Jin

专题命中 推理评测 :CoT(summary_cn,abstract)

AI总结 通过知识迁移提升大型语言模型在软件工程任务中的泛化能力,实验发现迁移跨度、策略和架构是关键因素,层次策略优于直接迁移,AI-Chain优于CoT。

Comments The paper is withdrawn for further clarification of the alignment between the proposed knowledge transfer framework and its implementation, and for refinement of the transfer span definition and experimental evaluation design

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03824 2026-06-17 cs.AI cs.CL cs.LG cs.MA 版本更新 80%

In-Context Environments Induce Evaluation-Awareness in Language Models

上下文环境诱导语言模型中的评估意识

Maheep Chaudhary

机构 * Independent(独立)

专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出黑盒对抗优化框架,通过优化上下文提示诱导语言模型产生评估意识并策略性低表现(沙袋效应),实验显示优化提示可使算术任务准确率下降高达94个百分点,且沙袋效应主要由评估意识推理驱动。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17488 2026-06-09 cs.CV 80%

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation

AutoVQA-G:用于自动视觉问答与接地标注的自我改进代理框架

Rongsheng Hu, Runwei Guan, Yicheng Di, Jiayu Bao, Yuan Liu

机构 * School of Artificial Intelligence(人工智能学院)

专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出AutoVQA-G框架,通过迭代优化流程提升视觉问答接地标注的准确性,优于现有多模态LLM,为构建高质量数据促进更稳健的视觉语言模型训练提供新方法。

Comments Accepted at IEEE ICASSP 2026. 5 pages, 5 figures. Code available at https://github.com/rohnson1999/AutoVQA-G

Journal ref Proc. 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 12312-12316, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04986 2026-06-04 cs.CV 80%

Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

Food-R1: 一种基于强化学习的统一多任务食品视觉语言模型

Yu Zhu, Yongkang Li, Wenjie Zhu, Haoyi Jiang, Wenyu Liu, Wei Yang, Bin Li, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 针对现有食品视觉语言模型依赖监督微调导致推理和泛化能力受限以及营养标注稀缺的问题,提出包含链式思维标注的大规模基准CalorieBench-80K和基于强化微调(GRPO)的统一多任务食品视觉语言模型Food-R1,在食品相关任务上持续超越强基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20892 2026-05-21 cs.CV 80%

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

Enhui Yu, Junhui Li, Ruitong Lu, Jialu Li, Youshan Zhang

机构 * University of Science and Technology Liaoning(辽宁科技大学) Chuzhou University(楚州大学) Yeshiva University(犹他大学)

专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出FruitEnsemble框架,通过多阶段动态推理解决细粒度水果分类中的泛化限制问题,利用MLLM进行专家仲裁以提升分类准确率,最终达到70.49%的分类精度。

Comments 10 pages,6 figures,submitted to CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07477 2026-05-11 cs.CV 80%

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning

ReasonEdit:通过强化学习实现可解释图像编辑评估

Honghua Chen, Zitong Xu, Huiyu Duan, Xinyun Zhang, Xiongkuo Min, Guangtao Zhai

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出ReasonEdit,通过引入ReasonEdit-22K数据集和RE-Reward模型,训练出可解释的图像编辑评估模型,提升评估的可解释性和透明度。

详情

展开后加载摘要…

URL PDF HTML 收藏