arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1129 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1129 篇

2407.17227 2024-07-25 cs.AI cs.CL 62%

LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover

Zijian Wu, Jiayu Wang, Dahua Lin, Kai Chen

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00744 2024-07-02 cs.AI cs.LG q-bio.NC 62%

Disentangled Representations for Causal Cognition

Filippo Torresan, Manuel Baltieri

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 49 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12147 2024-06-12 cs.AI cs.CL 62%

Eliciting Problem Specifications via Large Language Models

Robert E. Wray, James R. Kirk, John E. Laird

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 18 pages, Appendix. Revised in response to reviewer feedback. Accepted for Advances in Cognitive Systems (Jun 2024, Palermo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09140 2024-06-03 cs.CL cs.AI cs.IR 62%

Multi-hop Question Answering

Vaibhav Mavi, Anubhav Jangra, Adam Jatowt

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Published at Foundations and Trends in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05209 2024-04-03 cs.AI cs.CL 62%

HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models

Navapat Nananukul, Mayank Kejriwal

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted and orally presented in "SPIE Defense + Commercial Sensing (DCS 2024)" in National Harbor, Maryland, April 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18239 2024-04-02 cs.AI cs.CL cs.FL cs.RO 62%

Fine-Tuning Language Models Using Formal Methods Feedback

Yunhao Yang, Neel P. Bhatt, Tyler Ingebrand, William Ward, Steven Carr, Zhangyang Wang, Ufuk Topcu

专题命中 代码与定理证明 :planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04488 2024-03-19 cs.LG cs.AI cs.LO 62%

Magnushammer: A Transformer-Based Approach to Premise Selection

Maciej Mikuła, Szymon Tworkowski, Szymon Antoniak, Bartosz Piotrowski, Albert Qiaochu Jiang, Jin Peng Zhou, Christian Szegedy, Łukasz Kuciński, Piotr Miłoś, Yuhuai Wu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10051 2024-02-16 cs.AI cs.CL 62%

SwissNYF: Tool Grounded LLM Agents for Black Box Setting

Somnath Sendhil Kumar, Dhruv Jain, Eshaan Agarwal, Raunak Pandey

专题命中 代码与定理证明 :planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06973 2023-12-13 cs.AI cs.LG cs.LO 62%

Anytime Approximate Formal Feature Attribution

Jinqiang Yu, Graham Farr, Alexey Ignatiev, Peter J. Stuckey

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01866 2023-11-06 cs.CL cs.AI 62%

Towards Concept-Aware Large Language Models

Chen Shani, Jilles Vreeken, Dafna Shahaf

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2023 findings long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14250 2023-10-31 cs.CL cs.AI 62%

Language Models with Rationality

Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson, Hinrich Schuetze, Peter Clark

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05689 2023-09-13 cs.CL cs.AI 62%

Large Language Model for Science: A Study on P vs. NP

Qingxiu Dong, Li Dong, Ke Xu, Guangyan Zhou, Yaru Hao, Zhifang Sui, Furu Wei

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 73 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04505 2023-06-08 cs.LG cs.AI cs.CC cs.CR 62%

Hardness of Deceptive Certificate Selection

Stephan Wäldchen

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07961 2023-05-18 cs.IR cs.CL cs.LG 62%

Leveraging Large Language Models in Conversational Recommender Systems

Luke Friedman, Sameer Ahuja, David Allen, Zhenning Tan, Hakim Sidahmed, Changbo Long, Jun Xie, Gabriel Schubiner, Ajay Patel, Harsh Lara, Brian Chu, Zexi Chen, Manoj Tiwari

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12283 2023-02-21 cs.AI cs.LG 62%

Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs

Albert Q. Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothée Lacroix, Yuhuai Wu, Guillaume Lample

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04901 2022-11-15 cs.CL cs.LG 62%

Exploring Length Generalization in Large Language Models

Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, Behnam Neyshabur

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12910 2022-11-02 cs.CL cs.AI 62%

NaturalProver: Grounded Mathematical Proof Generation with Language Models

Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, Yejin Choi

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10914 2022-08-24 cs.LG cs.AI 62%

Home Run: Finding Your Way Home by Imagining Trajectories

Daria de Tinguy, Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI、cs.LG

Comments accepted for International Workshop on Active Inference IWAI 2022, ECML PKDD workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.00558 2021-10-04 cs.CL cs.AI cs.LO 62%

Natural language understanding for logical games

Adrian Groza, Cristian Nitu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.00739 2021-04-05 cs.SE cs.AI cs.LG cs.PL q-bio.OT 62%

Formal Methods for the Informal Engineer: Workshop Recommendations

Gopal Sarma, James Koppel, Gregory Malecha, Patrick Schultz, Eric Drexler, Ramana Kumar, Cody Roux, Philip Zucker

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI、cs.LG

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02231 2021-01-08 cs.AI cs.LG 62%

Controlling Synthetic Characters in Simulations: A Case for Cognitive Architectures and Sigma

Volkan Ustun, Paul S. Rosenbloom, Seyed Sajjadi, Jeremy Nuttal

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

Journal ref Interservice/Industry Training, Simulation, and Education Conference (I/ITSEC) 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02185 2021-01-07 cs.AI cs.LG 62%

Adaptive Synthetic Characters for Military Training

Volkan Ustun, Rajay Kumar, Adam Reilly, Seyed Sajjadi, Andrew Miller

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

Journal ref 2020 Interservice/Industry Training, Simulation, and Education Conference (I/ITSEC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05867 2020-05-06 cs.CL cs.AI 62%

Transformers as Soft Reasoners over Language

Peter Clark, Oyvind Tafjord, Kyle Richardson

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments IJCAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.00840 2019-09-19 cs.AI cs.CL 62%

Knowledge Authoring and Question Answering with KALM

Tiantian Gao

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments In Proceedings ICLP 2019, arXiv:1909.07646

Journal ref EPTCS 306, 2019, pp. 389-395

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.01458 2019-09-12 cs.MA cs.AI cs.LG 62%

Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning

Jakob N. Foerster, Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, Michael Bowling

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1309.4962 2013-09-20 cs.AI cs.DL cs.LG cs.LO cs.MS 62%

HOL(y)Hammer: Online ATP Service for HOL Light

Cezary Kaliszyk, Josef Urban

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18404 2026-04-21 cs.AI 61%

Six Llamas: Comparative Religious Ethics Through LoRA-Adapted Language Models

六头 llama:通过 LoRA 调整的语言模型进行比较宗教伦理研究

Chad Coleman, W. Russell Neuman, Manan Shah, Ali Dasdan, Matthew Crispi, Morris Chiang, Zack Leitman, Mustafa Poonawala

机构 * Meta

专题命中 代码与定理证明 :reasoning(abstract,comments);分类 cs.AI

AI总结 本文通过 LoRA 调整的语言模型比较不同宗教文本对伦理推理的影响,发现调整模型在伦理推理上与基础模型有系统性差异,并在不同温度设置下表现出稳定性与多样性。

Comments 51 pages, 14 figures. We present Six Llamas, a comparative study examining whether Llama-3.1-8B models fine-tuned on distinct religious corpora encode systematically different patterns of ethical reasoning. Five LoRA-adapted variants are constructed for Christianity, Islam, Judaism, Hinduism, and Buddhism. For theoretical background on the condensate comparative method, see arXiv:2603.07329

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10249 2025-09-15 cs.AI 61%

Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering

Hanna Abi Akl

机构 * Université Côte d’Azur(法国滨海大学) Inria(法国国家信息与自动化技术研究所) CNRS(法国国家科学研究中心) I3S(国际科学计算研究所) Sophia Antipolis, France(法国索菲亚大学) Data ScienceTech Institute (DSTI)(数据科学技术研究所)

专题命中 代码与定理证明 :reasoning(abstract,comments);分类 cs.AI

Comments accepted for the International Joint Conference on Rules and Reasoning (RuleML+RR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17127 2026-08-19 cs.CL 版本更新 57%

The Emergence of Lab-Driven Alignment Signatures: A Psychometric Framework for Auditing Latent Bias and Compounding Risk in Generative AI

实验室驱动的对齐签名的出现:一种心理测量框架,用于审计生成AI中的潜在偏见和叠加风险

Dusan Bosnjakovic

机构 * AI Researcher(人工智能研究员)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

AI总结 本文提出了一种心理测量框架,用于审计生成AI中的潜在偏见和叠加风险,通过分析九个领先模型的实验室信号,揭示了持续行为聚类的成因。

Comments v2: expanded from 9 to 18 behavioral dimensions and from 4 to 6 developer organizations; revised statistical methodology (rank-based inference with effect-size criterion, replacing variance-decomposition approach); model-level results now reported; references corrected throughout

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10818 2026-08-19 cs.AI cs.HC 版本更新 57%

LLM Enhancement with Domain Expert Mental Model to Reduce LLM Hallucination with Causal Prompt Engineering

结合领域专家心智模型的大语言模型增强:通过因果提示工程减少大语言模型幻觉

Boris Kovalerchuk, Brent D. Fegley

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

AI总结 本文提出结合专家心智模型(EMM)的因果提示工程框架,通过形式化相关前置流程构建EMM,减少LLM幻觉,在三类任务中验证了方法有效性。

Comments 42 pages,4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏