arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-02-18 至 2026-02-18 共收录 7 信号源:cs.CL, cs.AI, cs.LG

1. 逻辑推理 7 篇

2504.06438 2026-02-18 cs.CL cs.AI 85%

Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning

不要让它幻觉:通过检索增强的逻辑推理进行前提验证

Yuehan Qin, Shawn Li, Yi Nian, Xinyan Velocity Yu, Yue Zhao, Xuezhe Ma

机构 * University of Southern California(南加州大学)

专题命中 逻辑推理 :reasoning(title);logical reasoning(title);分类 cs.CL、cs.AI

AI总结 本文提出一种基于检索增强的逻辑推理方法,用于在生成前验证用户查询中的前提,从而减少幻觉并提高事实准确性。

Comments TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15336 2026-02-18 cs.AR 71%

Human-AI Interaction: Evaluating LLM Reasoning on Digital Logic Circuit included Graph Problems, in terms of creativity in design and analysis

人类-人工智能交互:评估LLM在包含图问题的数字逻辑电路中的推理能力,就设计和分析的创造性而言

Yogeswar Reddy Thota, Setareh Rafatirad, Homayoun Houman, Tooraj Nikoubin

专题命中 逻辑推理 :reasoning(title)

AI总结 评估LLM在数字逻辑电路问题中的推理能力,发现其在复杂问题上与官方答案存在差距,且易陷入教科书模板。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22211 2026-02-18 cs.CL cs.AI 62%

LogiPart: Local Large Language Models for Data Exploration at Scale with Logical Partitioning

LogiPart: 用于大规模数据探索的本地大语言模型与逻辑分区

Tiago Fernandes Tavares

机构 * Tiago Fernandes Tavares(独立研究者)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 LogiPart通过本地大语言模型与逻辑分区技术,在大规模数据探索中实现高效、可解释的税目发现,克服传统方法的效率与成本限制。

Comments This version introduces a major architectural shift to Local LLMs and NLI-based assignment, scaling the framework to O(1) generative complexity. Formerly titled 'Question-Driven Analysis and Synthesis'

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15183 2026-02-18 cs.LG cs.CL 62%

Seeing to Generalize: How Visual Data Corrects Binding Shortcuts

看见以泛化:视觉数据如何纠正绑定捷径

Nicolas Buzeta, Felipe del Rio, Cristian Hinostroza, Denis Parra, Hans Lobel, Rodrigo Toro Icarte

机构 * Department of Computer Science, Pontificia Universidad Católica, Santiago, Chile(计算机科学系,天主教大学,圣地亚哥,智利)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 视觉数据训练可增强模型在单模态任务中的推理与泛化能力,通过改变绑定策略提升分布外性能。

Comments Submitted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15311 2026-02-18 cs.AI 57%

Aeon: High-Performance Neuro-Symbolic Memory Management for Long-Horizon LLM Agents

Aeon: 高性能神经符号记忆管理用于长时域大语言模型智能体

Mustafa Arslan

机构 * Independent Researcher(独立研究者)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI

AI总结 Aeon通过神经符号认知操作系统优化长时域大语言模型的记忆管理,实现高效内存压缩和加速,提升推理性能与稳定性。

Comments v3: Production hardening. Added INT8 quantization (5.6x dot product speedup, 3.1x compression), crash recovery via decoupled WAL (<1% overhead), unlimited text storage via sidecar blob arena with generational GC, and epoch-based reclamation for lock-free reads (P99 750ns under 16-thread contention). Revised for systems engineering clarity

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15294 2026-02-18 cs.AI 57%

EAA: Automating materials characterization with vision language model agents

EAA: 用视觉语言模型代理自动化材料表征

Ming Du, Yanqi Luo, Srutarshi Banerjee, Michael Wojcik, Jelena Popovic, Mathew J. Cherukara

机构 * Argonne National Laboratory(阿贡国家实验室) Advanced Photon Source(先进光子源) Data Science and Learning Division(数据科学与学习 division) Department of Radiation Oncology(放射肿瘤学系) Northwestern University(西北大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI

AI总结 EAA利用视觉语言模型代理自动化材料表征,通过多模态推理和工具增强操作提升显微镜实验效率,减少操作负担并降低用户专业门槛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15228 2026-02-18 cs.SE 50%

An Empirical Study on the Effects of System Prompts in Instruction-Tuned Models for Code Generation

对指令微调模型在代码生成中系统提示效果的实证研究

Zaiyu Cheng, Antonio Mastropaolo

专题命中 逻辑推理 :reasoning(abstract)

AI总结 本研究探讨了系统提示对代码生成模型性能的影响,发现提示具体性与模型规模、编程语言等因素相关,且不同语言对提示敏感性存在差异。

Comments 34 pages, 12 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏