arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1129 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1129 篇

2512.06178 2025-12-09 cs.SE 50%

Systematically Thinking about the Complexity of Code Structuring Exercises at Introductory Level

在初级课程中系统思考代码结构练习的复杂性

Georgiana Haldeman, Peter Ohmann, Paul Denny

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本文提出一个框架,用于系统评估初级课程中代码结构练习的复杂性,通过定义重复、代码模式和数据依赖性三个维度,帮助设计提升学生分解与抽象能力的教育任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03635 2025-12-04 cs.LO cs.SE 50%

Formal Analysis of the Sigmoid Function and Formal Proof of the Universal Approximation Theorem

对Sigmoid函数的正式分析及通用逼近定理的正式证明

Dustin Bryant, Jim Woodcock, Simon Foster

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本文在Isabelle/HOL中形式化分析了Sigmoid函数并证明了通用逼近定理,填补了形式化证明库的空白,提升了神经网络的可信度。

Comments 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02544 2025-12-03 cs.CY 50%

A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty

以人为中心的AI意识伦理争论框架

Zhou Ziheng, Haiqiang Dai, Bin Ling, Ying Nian Wu, Demetri Terzopoulos

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本文提出以人为中心的AI意识伦理框架,通过哲学不确定性的三级结构,确立无意识假设、风险谨慎和透明推理原则,以平衡哲学严谨性与实践指导,为AI伦理问题提供负责任的演进路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23188 2025-12-01 cs.HC 50%

Can Intelligent User Interfaces Engage in Philosophical Discussions? A Longitudinal Study of Philosophers' Evolving Perceptions

智能用户界面能否参与哲学讨论?哲学家们观念演变的纵向研究

Yibo Meng, Lyumanshan Ye, Eve He, Zhe Yan, Zhiming Liu, Yipeng Yu, Yan Guan, Xiaolan Ding

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本研究通过纵向研究探讨哲学学者对生成式AI智能用户界面参与哲学讨论态度的演变,发现其从抵制到接受再到质疑的三阶段变化,揭示IUI在哲学推理中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18509 2025-11-27 cs.PL 50%

Higher-Order Behavioural Conformances via Fibrations

通过纤维化实现更高阶的行为一致性

Henning Urbat

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本文提出了一种统一的范畴方法,通过纤维化实现高阶语言的行为一致性,证明了最大行为一致性在一般条件下的合同性质。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18438 2025-11-25 cs.CR cs.SE 50%

LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework

LLMs作为固件专家:一种运行时增长的代理树框架

Xiangrui Zhang, Zeyu Chen, Haining Wang, Qiang Li

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本文提出FIRMHIVE框架,利用LLMs作为固件安全分析师,通过递归代理蜂群实现更深入的固件漏洞检测,提升漏洞识别率和精度。

Comments 18 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22085 2025-11-25 cs.SE 50%

BOOP: Write Right Code

BOOP:写出正确的代码

Vaani Goenka, Aalok Thakkar

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 BOOP通过结构化框架提升编程教育中的计算思维能力,强调正确性优先的开发方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09971 2025-11-14 cs.CR 50%

Proofs of Useful Work from Arbitrary Matrix Multiplication

Ilan Komargodski, Omri Weinstein

专题命中 代码与定理证明 :verifier(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15112 2025-11-05 cs.CR 50%

AndroByte: LLM-Driven Privacy Analysis through Bytecode Summarization and Dynamic Dataflow Call Graph Generation

Mst Eshita Khatun, Lamine Noureddine, Zhiyong Sui, Aisha Ali-Gombe

专题命中 代码与定理证明 :reasoning(abstract)

Comments Accepted at the Annual Computer Security Applications Conference (ACSAC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17657 2025-10-24 cs.DB 50%

My Ontologist: Evaluating BFO-Based AI for Definition Support

Carter Benson, Alec Sculley, Austin Liebers, John Beverley

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01536 2025-10-22 cs.IR 50%

AI4DiTraRe: Building the BFO-Compliant Chemotion Knowledge Graph

Ebrahim Norouzi, Nicole Jung, Anna M. Jacyszyn, Jörg Waitelonis, Harald Sack

专题命中 代码与定理证明 :reasoning(abstract)

Comments 12 pages, 7 figures. Camera-ready version. Accepted to the 5th International Workshop on Scientific Knowledge: Representation, Discovery, and Assessment; 2 November 2025 - Nara, Japan; co-located with The 24th International Semantic Web Conference, ISWC 2025. Published in CEUR proceedings Vol-4065, pages 45-56

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09585 2025-10-20 cs.CV 50%

Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation

Jitesh Jain, Zhengyuan Yang, Humphrey Shi, Jianfeng Gao, Jianwei Yang

机构 * Microsoft Research, Redmond(微软研究院(红mond)) Meta Superintelligence Labs(Meta超智能实验室)

专题命中 代码与定理证明 :reasoning(abstract)

Comments Project Page: https://praeclarumjj3.github.io/visper_lm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10148 2025-10-14 cs.SE 50%

A Systematic Study on Generating Web Vulnerability Proof-of-Concepts Using Large Language Models

Mengyao Zhao, Kaixuan Li, Lyuye Zhang, Wenjing Dang, Chenggong Ding, Sen Chen, Zheli Liu

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03874 2025-10-08 quant-ph 50%

Any theory that admits a Wigner's Friend type multi-agent paradox is logically contextual

Nuriya Nurgalieva, V. Vilasini

专题命中 代码与定理证明 :reasoning(abstract)

Comments 39+16 pages. Both authors contributed equally to this work. Initial versions of some of these results were included in NN's PhD thesis (ETH Zurich, 2023). v2 includes additional citations and clarifications in the outlook section

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04689 2025-10-07 cs.SE 50%

Evolaris: A Roadmap to Self-Evolving Software Intelligence Management

Chengwei Liu, Wenbo Guo, Yuxin Zhang, Limin Wang, Sen Chen, Lei Bu, Yang Liu

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23675 2025-09-30 cs.SE 50%

PAT-Agent: Autoformalization for Model Checking

Xinyue Zuo, Yifan Zhang, Hongshu Wang, Yufan Cai, Zhe Hou, Jing Sun, Jin Song Dong

专题命中 代码与定理证明 :planning(abstract)

Comments Accepted in ASE 2025 (International Conference on Automated Software Engineering)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19854 2025-09-25 cs.LO math.AC 50%

L-Mosaics and Bounded Join-Semilattices in Isabelle/HOL

Alessandro Linzi

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15754 2025-09-25 cs.CR cs.PL cs.SE 50%

Hornet Node and the Hornet DSL: A Minimal, Executable Specification for Bitcoin Consensus

Toby Sharp

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02878 2025-09-04 cs.HC 50%

Designing a Lightweight GenAI Interface for Visual Data Analysis

Ratanond Koonchanok, Alex Kale, Khairi Reda

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18675 2025-08-27 cs.SE 50%

Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision

Xu Lu, Weisong Sun, Yiran Zhang, Ming Hu, Cong Tian, Zhi Jin, Yang Liu

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06342 2025-08-11 cs.CV cs.SI 50%

Street View Sociability: Interpretable Analysis of Urban Social Behavior Across 15 Cities

Kieran Elrod, Katherine Flanigan, Mario Bergés

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 代码与定理证明 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02733 2025-08-06 cs.SE cs.HC 50%

What's in a Proof? Analyzing Expert Proof-Writing Processes in F* and Verus

Rijul Jain, Shraddha Barke, Gabriel Ebner, Md Rakib Hossain Misu, Shan Lu, Sarah Fakhoury

专题命中 代码与定理证明 :verifier(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14807 2025-08-05 astro-ph.IM astro-ph.GA 50%

Interpreting Multi-band Galaxy Observations with Large Language Model-Based Agents

Zechang Sun, Yuan-Sen Ting, Yaobo Liang, Nan Duan, Song Huang, Zheng Cai

专题命中 代码与定理证明 :reasoning(abstract)

Comments Accepted at the NIPS ML4PS Workshop 2024. The journal version is in preparation. Code and data will be fully made public following the journal publication. We welcome any comments and feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15241 2025-07-22 cs.SE 50%

FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents

Vikram Nitin, Baishakhi Ray, Roshanak Zilouchian Moghaddam

专题命中 代码与定理证明 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04857 2025-07-08 cs.SE 50%

Supporting Software Formal Verification with Large Language Models: An Experimental Study

Weiqi Wang, Marie Farrell, Lucas C. Cordeiro, Liping Zhao

专题命中 代码与定理证明 :verifier(abstract)

Comments Accepted for publication in 2025 IEEE 33rd International Requirements Engineering Conference (RE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18796 2025-06-24 cs.SE 50%

Context-Aware CodeLLM Eviction for AI-assisted Coding

Kishanthan Thangarajah, Boyuan Chen, Shi Chang, Ahmed E. Hassan

专题命中 代码与定理证明 :reasoning(abstract)

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16639 2025-06-23 cs.SE 50%

LLM-based Satisfiability Checking of String Requirements by Consistent Data and Checker Generation

Boqi Chen, Aren A. Babikian, Shuzhao Feng, Dániel Varró, Gunter Mussbacher

专题命中 代码与定理证明 :reasoning(abstract)

Comments Accepted at the 33rd IEEE International Requirements Engineering 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01536 2025-06-04 quant-ph 50%

Quantum Agents

Eldar Sultanow, Madjid Tehrani, Siddhant Dutta, William J Buchanan, Muhammad Shahbaz Khan

专题命中 代码与定理证明 :planning(abstract)

Comments 45 Pages, 16 figures, 3 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23311 2025-05-30 cs.LO cs.AR cs.SC 50%

Towards LLM-based Generation of Human-Readable Proofs in Polynomial Formal Verification

Rolf Drechsler

专题命中 代码与定理证明 :reasoning(abstract)

Comments 4 pages; keynote given at 7th International Symposium on Devices, Circuits and Systems (ISDCS 2025), May 27-30, 2025, IIEST Shibpur, Kolkata, India

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00511 2025-04-22 math.OC cs.SY eess.SY math.CT 50%

A Bayesian Interpretation of the Internal Model Principle

Manuel Baltieri, Martin Biehl, Matteo Capucci, Nathaniel Virgo

专题命中 代码与定理证明 :reasoning(abstract)

Comments 14 pages, no figures

详情

展开后加载摘要…

URL PDF HTML 收藏