Finding the most similar textual documents using Case-Based Reasoning
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2019
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments 10 pages, 7 tables, 2 figures
MindClaw: 用于精确干预的闭环具身心理状态推理
机构 * Jilin University(吉林大学) ; Microsoft Asia(微软亚洲) ; National Taiwan University(国立台湾大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 提出MindClaw框架,通过闭环具身心理状态推理实现精确干预,结合多源输入、信念记忆、认知触发技能和动作生成,在动态环境中优化干预时机。
Comments Extended version of the CVPR 2026 paper *MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents*. This work is in progress
S-AI-Recursive:一种生物启发式且时间稀疏的AI架构,用于迭代、反思和节能推理
机构 * Mohammed V University(穆莱·伊斯梅尔大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 本文提出S-AI-Recursive架构,通过激素闭环迭代而非单向传递实现推理,结合生物启发式方法和数学模型,验证了时间稀疏性原理。
Comments Preprint. 55 pages. No figures. S-AI-Recursive: Convergent Recursive Reasoning
为策略制定者搭建框架:霍特林空间市场中与架构相关的推理干预
机构 * California Institute of Technology(加州理工学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 研究结构化推理干预对大语言模型战略经济推理的影响及与模型架构的关系,以霍特林模型评估GPT - 4.1 - mini和GPT - 5 - mini,发现支架类型与模型架构有交叉交互作用,对抗性测试有损害,还存在陈述性 - 程序性差距。
Comments 26 pages (11 main + 15 appendix), 6 figures, 4 tables. Accepted at the ICLR 2026 Workshop on LLM Reasoning
面向旅游的推理大语言模型:基于领域特定知识图谱
机构 * Expedia Group(Expedia集团)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
AI总结 提出一种模块化流水线,利用专家构建的知识图谱生成多跳问答对,微调大语言模型以提升旅游领域推理的准确性和校准能力,在基准上达到82.4%精确匹配。
Comments Accepted to the Uncertainty Reasoning and Quantification in Decision Making (UDM) Workshop, KDD 2026 (To be presented in August 2026)
ComBench: 奥林匹克级组合数学中严格证明推理与构造实现的基准测试
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Peking University(北京大学) ; Shanghai Jiao Tong University(上海交通大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 提出ComBench基准,包含100道奥林匹克级组合问题,分分析和构造两类,通过评分与验证评估大模型推理能力,发现最强模型准确率仅65.4%,且证明推理与构造实现能力存在差异。
Comments 39 pages, 6 figures, 26 tables. Project page: https://simplified-reasoning.github.io/ComBench/docs/
视觉-表格问答:面向表格图像推理的开放领域基准
机构 * École de Technologie Supérieure(埃克塞技术学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
AI总结 本文提出Visual-TableQA,一个大规模开放领域多模态数据集,用于评估和提升视觉推理能力,包含2500个LaTeX渲染表格和6000对推理密集型问答对,通过多模型协作生成,验证了模型在外部基准上的鲁棒性。
Comments Accepted at the First Workshop on Foundations of Reasoning in Language Models, NeurIPS 2025. Available at: https://openreview.net/forum?id=fvJRsGwhPf
基于LLM的多模态推理用于加密流量解释:一个基准
机构 * School of Microelectronics and Communication Engineering, Chongqing University(重庆大学微电子与通信工程学院) ; School of Data Science, Lingnan University(岭南大学数据科学学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 本文提出BGTD基准和mmTraffic框架,通过结合原始字节与结构化注释,实现可解释的加密流量解释,生成高保真的人可读报告,同时保持高分类准确率。
Comments Project page \url{https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark}
并非其他:一种区分推理与记忆的通用技术,用于多选LLM评估基准
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
AI总结 本文提出一种通用方法,通过改变数学问题的数值来区分LLM的推理能力与记忆能力,评估了多个模型在公开和私有数据集上的表现,发现模型在该方法下准确率显著下降,揭示了记忆在当前LLM回答中的重要作用。
Journal ref "On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks," in IEEE Access, vol. 14, pp. 9384-9393, 2026
HeaRT:一种基于分层电路推理树的代理框架用于AMS设计优化
机构 * ECE Department, The University of Texas at Austin(德克萨斯大学奥斯汀分校电子与计算机工程系) ; NVIDIA Corporation(英伟达公司) ; The George Washington University(乔治华盛顿大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 HeaRT提出了一种分层电路推理树的代理框架,通过提升F1(subcircuits)和F1(loops)指标,实现更高效的AMS设计优化,且在不同架构上表现出更好的适应性和收敛速度。
Comments Analog Design Automation, Hierarchical Circuit Reasoning, Context-Aware Design Adaptation, LLMs, Agentic Frameworks, Electronic Design Automation (EDA)
小模型和推理大模型能否对期刊文章进行科研质量评分?平均和少样本学习是否有帮助?
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 本文评估了小模型和推理模型对期刊文章科研质量评分的能力,发现4b以上的小模型在使用评分平均时表现良好,但推理模型无明显优势。
Comments Thelwall, M. & Mohammadi, E. (2026). Can small and reasoning Large Language Models score journal articles for research quality and do averaging and few-shot help? Scientometrics
基于像素级精度的推理:QVLM架构与SQuID数据集用于定量遥感分析
机构 * Department of Machine Learning, NEC Laboratories America(机器学习系,NEC美国实验室)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 本文提出QVLM架构和SQuID数据集,通过解耦语言理解和视觉分析,提升定量空间推理的准确性。
Comments Submitted to CVPR 2026. Introduces the QVLM architecture and the SQuID dataset for quantitative geospatial reasoning. Dataset DOI: 10.57967/hf/7565
机构 * Harvard University(哈佛大学) ; MIT Media Lab(麻省理工学院媒体实验室)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
Comments Accepted at the Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2025
机构 * Tencent Youtu Lab(腾讯优图实验室)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
Comments Accepted to NeurIPS 2025 Workshop on Efficient Reasoning
机构 * Criteo AI Lab(Criteo人工智能实验室) ; Ecole Polytechnique Paris(巴黎高等理工学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG
Comments Accepted to the Efficient Reasoning Workshop at NeuRIPS 2025
机构 * School of Computer Science, Australian National University and CSIRO(计算机科学学院,澳大利亚国立大学和CSIRO)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
Comments The original version of this article was withdrawn because there were errors in the evaluation of model faithfulness to reasoning strategies and completeness of reasoning. The analysis was re-conducted correctly and version two contains the corrections
机构 * Tsinghua University(清华大学) ; OpenDataLab, Shanghai Artificial Intelligence Laboratory(开放数据实验室、上海人工智能实验室) ; Renmin University of China(中国人民大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
Comments REST (Reasoning Evaluation through Simultaneous Testing), a stress-testing framework that concurrently exposes LRMs to multiple problems simultaneously
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
Comments Proceedings of the 1st Workshop on Natural Language Reasoning and Structured Explanations (NLRSE) at ACL 2023
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
Comments 6 pages, 1 figure, 2 tables, SemEval -2020, Commonsense Reasoning and Natural Language Processing
基于智能体的多视角聚合测试断言生成
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)
AI总结 针对现有LLM断言生成方法的局限,提出基于智能体的AssertMate框架,通过三个组件聚合多视角,在Defects4J和EvoSuite验证中性能显著优于现有技术。
自适应智能释放:大型语言模型中知识迁移的可行性
专题命中 推理评测 :CoT(summary_cn,abstract)
AI总结 通过知识迁移提升大型语言模型在软件工程任务中的泛化能力,实验发现迁移跨度、策略和架构是关键因素,层次策略优于直接迁移,AI-Chain优于CoT。
Comments The paper is withdrawn for further clarification of the alignment between the proposed knowledge transfer framework and its implementation, and for refinement of the transfer span definition and experimental evaluation design
上下文环境诱导语言模型中的评估意识
机构 * Independent(独立)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出黑盒对抗优化框架,通过优化上下文提示诱导语言模型产生评估意识并策略性低表现(沙袋效应),实验显示优化提示可使算术任务准确率下降高达94个百分点,且沙袋效应主要由评估意识推理驱动。
AutoVQA-G:用于自动视觉问答与接地标注的自我改进代理框架
机构 * School of Artificial Intelligence(人工智能学院)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)
AI总结 本文提出AutoVQA-G框架,通过迭代优化流程提升视觉问答接地标注的准确性,优于现有多模态LLM,为构建高质量数据促进更稳健的视觉语言模型训练提供新方法。
Comments Accepted at IEEE ICASSP 2026. 5 pages, 5 figures. Code available at https://github.com/rohnson1999/AutoVQA-G
Journal ref Proc. 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 12312-12316, 2026
Food-R1: 一种基于强化学习的统一多任务食品视觉语言模型
机构 * Huazhong University of Science and Technology(华中科技大学)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)
AI总结 针对现有食品视觉语言模型依赖监督微调导致推理和泛化能力受限以及营养标注稀缺的问题,提出包含链式思维标注的大规模基准CalorieBench-80K和基于强化微调(GRPO)的统一多任务食品视觉语言模型Food-R1,在食品相关任务上持续超越强基线。
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
机构 * University of Science and Technology Liaoning(辽宁科技大学) ; Chuzhou University(楚州大学) ; Yeshiva University(犹他大学)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)
AI总结 本文提出FruitEnsemble框架,通过多阶段动态推理解决细粒度水果分类中的泛化限制问题,利用MLLM进行专家仲裁以提升分类准确率,最终达到70.49%的分类精度。
Comments 10 pages,6 figures,submitted to CVPR 2026
Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2026
ReasonEdit:通过强化学习实现可解释图像编辑评估
机构 * University of Electronic Science and Technology of China(电子科学与技术大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)
AI总结 本文提出ReasonEdit,通过引入ReasonEdit-22K数据集和RE-Reward模型,训练出可解释的图像编辑评估模型,提升评估的可解释性和透明度。