arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1129 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1129 篇

2601.15737 2026-01-23 cs.AI cs.CL 62%

PhysProver: Advancing Automatic Theorem Proving for Physics

PhysProver: 推动物理领域自动定理证明的发展

Hanning Zhang, Ruida Wang, Rui Pan, Wenyuan Wang, Bingxu Meng, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rutgers University(罗格斯大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 PhysProver通过结合可验证语言和强化学习,提升物理领域形式化定理证明的效率和效果。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09709 2026-01-22 cs.CL cs.LG 62%

Large Language Models Encode Semantics and Alignment in Linearly Separable Representations

大语言模型在线性可分的表示中编码语义和对齐

Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, Amin Karbasi

机构 * Yale University(耶鲁大学) Foundation AI – Cisco Systems Inc(Foundation AI – 卡西欧系统公司)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 本研究发现大语言模型通过线性可分的表示编码语义和对齐,提出基于潜在空间的MLP探测器有效提升安全防护。

Comments IJCNLP and the Asian Chapter of ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11622 2026-01-21 cs.AI cs.LG 62%

Dynamical Systems Analysis Reveals Functional Regimes in Large Language Models

动力系统分析揭示大语言模型中的功能模式

Hassan Ugail, Newton Howard

机构 * Centre for Visual Computing and Intelligent Systems(视觉计算与智能系统中心) University of Bradford(布拉德福德大学) School of Individualised Study(个性化学习学院) Rochester Institute of Technology(罗切斯特技术学院)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文通过动力系统分析揭示大语言模型在不同功能模式中的计算组织差异,提出一种基于时间序列的动态度量指标,用于评估模型内部动态特性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10077 2026-01-21 cs.NE cs.AI cs.DS cs.LG 62%

Predictive Spike Timing Enables Distributed Shortest Path Computation in Spiking Neural Networks

预测性脉冲时间使分布式最短路径计算在脉冲神经网络中成为可能

Simen Storesund, Kristian Valset Aars, Robin Dietrich, Nicolai Waniek

机构 * Department of Mathematical Sciences(数学科学系) Norwegian University of Science and Technology(挪威科学与技术大学) School of Computation, Information and Technology(计算、信息与技术学院) Technical University Munich(慕尼黑技术大学)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI、cs.LG

AI总结 该研究提出了一种基于生物合理机制的最短路径计算算法,利用脉冲时间巧合实现分布式计算,为生物和人工系统中的复杂问题解决提供新思路。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06478 2026-01-06 cs.LG cs.AI 62%

Anytime-Valid Answer Sufficiency Certificates for LLM Generation via Sequential Information Lift

通过顺序信息提升实现LLM生成的任何时间有效性答案充分性证书

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction & Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

AI总结 通过顺序信息提升实现LLM生成的任何时间有效性答案充分性证书,利用经验动态形式提升方法,在减少生成长度的同时保持delta级误差控制,并通过轻量级正确性门提升端任务正确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00816 2026-01-06 cs.AI cs.CR cs.LG 62%

MathLedger: A Verifiable Learning Substrate with Ledger-Attested Feedback

MathLedger: 一种具有账本证明反馈的可验证学习基础

Ismail Ahmad Abdullah

机构 * CNU(中国矿业大学)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

AI总结 MathLedger通过整合形式验证、密码学证明和学习动态,提供一种可验证学习的基础,实现可审计的机器认知系统。

Comments 14 pages, 1 figure, 2 tables, 2 appendices with full proofs. Documents v0.9.4-pilot-audit-hardened audit surface with fail-closed governance, canonical JSON hashing, and artifact classification. Phase I infrastructure validation; no capability claims

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17904 2025-12-16 cs.CR cs.AI cs.CL 62%

BreakFun: Jailbreaking LLMs via Schema Exploitation

BreakFun: 通过模式利用对LLM进行劫持

Amirkia Rafiei Oskooei, Mehmet S. Aktas

机构 * Department of Computer Engineering, Yildiz Technical University(计算机工程系,伊兹密尔技术大学)

专题命中 代码与定理证明 :chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 BreakFun通过利用LLM对结构模式的遵守能力,揭示了其在劫持攻击中的脆弱性,并提出对抗性提示解构作为缓解策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15809 2025-12-02 cs.CL cs.AI cs.DB 62%

Chain-of-Query: Unleashing the Power of LLMs in SQL-Aided Table Understanding via Multi-Agent Collaboration

链式查询:通过多智能体协作释放LLM在SQL辅助表格理解中的潜力

Songyuan Sui, Hongyi Liu, Serena Liu, Li Li, Soo-Hyun Choi, Rui Chen, Xia Hu

机构 * Rice University(里士大学) Samsung Electronics America(三星电子美国分公司) Warner Bros. Discovery(华纳兄弟发现)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 链式查询通过多智能体协作提升SQL辅助表格理解的准确性与有效性

Comments AACL 2025 Main Conference (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21050 2025-11-27 cs.LG cs.AI stat.ML 62%

Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs

打破安全性与能力的权衡:具有可验证奖励的强化学习在LLMs中维持安全护栏

Dongkyu Derek Cho, Huan Song, Arijit Ghosh Chowdhury, Haotian An, Yawei Wang, Rohit Thekkanal, Negin Sokhandan, Sharlina Keshava, Hannah Marlowe

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过可验证奖励的强化学习方法,在提升LLM推理能力的同时维持安全性,挑战了传统安全性与能力权衡的假设。

Comments AAAI-26 Workshop on Post-AI Formal Methods

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11770 2025-11-18 cs.AI cs.LG 62%

Learning to Refine: An Agentic RL Approach for Iterative SPARQL Query Construction

Floris Vossebeld, Shenghui Wang

机构 * Faculty of Electrical Engineering, Mathematics and Computer Science(电气工程、数学与计算机科学学院) University of Twente(特文特大学) Microsoft Netherlands(微软荷兰)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08274 2025-11-12 cs.AI cs.CL 62%

Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs

Anton Gusarov, Anastasia Volkova, Valentin Khrulkov, Andrey Kuznetsov, Evgenii Maslov, Ivan Oseledets

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Code to be released

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25427 2025-10-30 cs.CL cs.AI 62%

RLMEval: Evaluating Research-Level Neural Theorem Proving

Auguste Poiroux, Antoine Bosselut, Viktor Kunčak

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings. RLMEval benchmark released: https://github.com/augustepoiroux/RLMEval

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24115 2025-10-29 cs.AI cs.LG 62%

HistoLens: An Interactive XAI Toolkit for Verifying and Mitigating Flaws in Vision-Language Models for Histopathology

Sandeep Vissapragada, Vikrant Sahu, Gagan Raj Gupta, Vandita Singh

机构 * Indian Institute of Technology(印度理工学院) All India Institute of Medical Sciences(全印度医学科学研究所)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17873 2025-10-28 cs.CL cs.AI cs.CE 62%

MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback

Wanhao Liu, Zonglin Yang, Jue Wang, Lidong Bing, Di Zhang, Dongzhan Zhou, Yuqiang Li, Houqiang Li, Erik Cambria, Wanli Ouyang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) MiroMind

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21894 2025-10-28 cs.CL cs.AI 62%

Understanding Network Behaviors through Natural Language Question-Answering

Mingzhe Xing, Chang Tian, Jianan Zhang, Lichen Pan, Peipei Liu, Zhaoteng Yan, Yinliang Yue

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01346 2025-10-14 cs.AI cs.CL 62%

Aristotle: IMO-level Automated Theorem Proving

Tudor Achim, Alex Best, Alberto Bietti, Kevin Der, Mathïs Fédérico, Sergei Gukov, Daniel Halpern-Leistner, Kirsten Henningsgard, Yury Kudryashov, Alexander Meiburg, Martin Michelsen, Riley Patterson, Eric Rodriguez, Laura Scharff, Vikram Shanker, Vladmir Sicca, Hari Sowrirajan, Aidan Swope, Matyas Tamas, Vlad Tenev, Jonathan Thomm, Harold Williams, Lawrence Wu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14756 2025-10-10 cs.LG cs.AI 62%

LLINBO: Trustworthy LLM-in-the-Loop Bayesian Optimization

Chih-Yu Chang, Milad Azvar, Chinedum Okwudire, Raed Al Kontar

机构 * University of Michigan(密歇根大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01272 2025-10-03 cs.AI cs.LG 62%

Modeling Others' Minds as Code

Kunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques, Max Kleiman-Weiner

机构 * Department of Computer Science, University of Washington(华盛顿大学计算机科学系) Department of Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25343 2025-10-01 cs.AI cs.CL 62%

Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks

Yiming Wang, Rui Wang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09871 2025-09-22 cs.CL cs.AI 62%

Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case

Bastián González-Bustamante, Nando Verelst, Carla Cisternas

机构 * Universidad Diego Portales(迪亚戈·波特莱斯大学) Leiden University(莱顿大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Working paper: 18 pages, 4 tables, 2 figures

Journal ref Empiria Lab Method Series (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12233 2025-09-17 cs.CR cs.AI cs.ET cs.LG cs.NI 62%

Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analytics

Meryem Malak Dif, Mouhamed Amine Bouchiha, Abdelaziz Amara Korba, Yacine Ghamri-Doudane

机构 * L3i - La Rochelle University, La Rochelle, France(L3i - 拉罗谢尔大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 7 figures, Accepted at LCN'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06902 2025-09-09 cs.CL cs.CR cs.DB cs.LG 62%

Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification

Aivin V. Solatorio

专题命中 代码与定理证明 :verifier(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01716 2025-09-03 cs.AI cs.CL 62%

An LLM-enabled semantic-centric framework to consume privacy policies

Rui Zhao, Vladyslav Melnychuk, Jun Zhao, Jesse Wright, Nigel Shadbolt

机构 * University of Oxford(牛津大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14870 2025-08-29 cs.LO cs.AI cs.LG 62%

Application of AI to formal methods - an analysis of current trends

Sebastian Stock, Jannik Dunkelau, Atif Mashkoor

机构 * Institute of Software System Engineering(软件系统工程研究所) Johannes Kepler University Linz(约翰·凯撒大学林茨分校) Faculty of Mathematics and Natural Sciences(数学与自然科学学院) Heinrich Heine University Düsseldorf(海因里希·海涅大学杜塞尔多夫分校)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14275 2025-08-21 cs.CL cs.AI 62%

Disentangling concept semantics via multilingual averaging in Sparse Autoencoders

Cliff O'Reilly, Ernesto Jimenez-Ruiz, Tillman Weyde

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06017 2025-08-11 cs.SE cs.CL cs.LG 62%

Position: Intelligent Coding Systems Should Write Programs with Justifications

Xiangzhe Xu, Shiwei Feng, Zian Su, Chengpeng Wang, Xiangyu Zhang

机构 * Purdue University(普渡大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02584 2025-08-05 cs.CL cs.AI 62%

MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification

Ming Pok Ng, Junqi Jiang, Gabriel Freedman, Antonio Rago, Francesca Toni

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14722 2025-07-22 cs.LG cs.AI 62%

LeanTree: Accelerating White-Box Proof Search with Factorized States in Lean 4

Matěj Kripner, Michal Šustr, Milan Straka

机构 * Charles University, Faculty of Mathematics and Physics(查尔斯大学数学与物理系) Czech Technical University, Faculty of Electrical Engineering(捷克技术大学电气工程系)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02726 2025-07-04 cs.AI cs.LG 62%

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving

Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov, Anthony Bordg, Jun Wang, Haitham Bou Ammar

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Imperial College London(伦敦帝国学院) Huawei Lagrange Center(华为拉格朗日中心) UCL Centre for AI(大学学院人工智能中心)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17161 2025-06-19 cs.CL cs.LG cs.LO 62%

Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence

İlker Işık, Ramazan Gokberk Cinbis, Ebru Aydin Gol

机构 * Department of Computer Engineering, Middle East Technical University, Ankara, Turkey(中欧技术大学计算机工程系) Microsoft, İstanbul, Turkey(微软公司)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

Comments ICML 2025 Poster Paper, Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏