arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-04 至 2025-12-04 共收录 25 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 25 篇

2506.23888 2025-12-04 cs.CL cs.LG 92%

Advancing Multi-Step Mathematical Reasoning in Large Language Models through Multi-Layered Self-Reflection with Auto-Prompting

通过多层自反思与自动提示推进大语言模型的多步数学推理

André de Souza Loureiro, Jorge Valverde-Rebaza, Julieta Noguez, David Escarcega, Ricardo Marcacini

机构 * Luiz de Queiroz College of Agriculture University of São Paulo(圣保罗大学农业学院) School of Engineering and Sciences Tecnologico de Monterrey(蒙特雷技术学院工程与科学学院) Institute of Mathematics and Computer Sciences University of São Paulo(圣保罗大学数学与计算机科学研究所)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);prompting(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出MAPS框架,通过多层自反思与自动提示提升大语言模型的多步数学推理能力,实验证明其在多个基准测试中优于标准CoT并接近优化模型。

Comments Accepted for publication in: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2025). Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03568 2025-12-04 cs.HC 90%

Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough

合成认知 walkthrough:使大语言模型性能与人类认知 walkthrough 对齐

Ruican Zhong, David W. McDonald, Gary Hsieh

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 研究探讨了大型语言模型能否通过模拟人类行为来提升认知 walkthrough 的效率,并展示了其在可用性测试中的潜在应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03955 2025-12-04 cs.AI cs.ET 89%

Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol

基于大语言模型代理的规划与控制基准:带有模型上下文协议的积木世界

Niklas Jobs, Luis Miguel Vieira da Silva, Jayanth Somashekaraiah, Maximilian Weigand, David Kube, Felix Gehlhoff

机构 * Institute of Automation Technology, Helmut Schmidt University / University of the Federal Armed Forces Hamburg, Germany(自动化技术研究所,海姆·施密特大学/联邦武装部队大学汉堡,德国) Siemens AG, Nuremberg, Germany(西门子股份公司,纽伦堡,德国)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出一个基于大语言模型的规划与控制基准,通过积木世界问题和模型上下文协议,提供五个复杂度类别以评估不同代理架构的性能。

Comments This work has been submitted to IFAC for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03722 2025-12-04 cs.NI 89%

Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks

无线网络中增强型强化学习教程:大型语言模型

Lingyi Cai, Wenjie Fu, Yuxi Huang, Ruichen Zhang, Yinqiu Liu, Jiawen Kang, Zehui Xiong, Tao Jiang, Dusit Niyato, Xianbin Wang, Shiwen Mao, Xuemin Shen

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文介绍了利用大型语言模型增强无线网络强化学习的方法,涵盖分类框架、角色功能及应用场景,探讨未来发展方向。

Comments 30 pages, 12 figures, survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03421 2025-12-04 cs.SE 89%

Exploring the Potential and Limitations of Large Language Models for Novice Program Fault Localization

探索大型语言模型在初级程序员故障定位中的潜力与局限性

Hexiang Xu, Hengyuan Liu, Yonghao Wu, Xiaolan Kang, Xiang Chen, Yong Liu

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文评估了LLM在初级程序员故障定位中的性能,发现高级模型在准确性上表现优异,但存在推理过度和计算成本高等问题。

Comments The paper has been accepted for publication in The Journal of Systems & Software

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09060 2025-12-04 cs.DB 89%

GenRewrite: Query Rewriting via Large Language Models

GenRewrite: 通过大语言模型实现查询重写

Jie Liu, Barzan Mozafari

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 GenRewrite 利用大语言模型实现查询重写,通过自然语言重写规则和反例引导技术提升查询优化效率。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03111 2025-12-04 cs.CR 85%

Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents

Les Dissonances:池中工具赋能LLM代理中的跨工具采集与污染

Zichuan Li, Jian Cui, Xiaojing Liao, Luyi Xing

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出了一种新型威胁XTHP,揭示了LLM代理中多工具整合带来的安全风险,并通过Chord工具验证了75%的现实世界工具存在此漏洞。

Comments Network and Distributed System Security (NDSS) Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03272 2025-12-04 cs.AI 84%

When Do Symbolic Solvers Enhance Reasoning in Large Language Models?

符号求解器在大语言模型中增强推理何时有效?

Zhiyuan He, Dingmin Wang

机构 * University College London(伦敦大学学院) University of Oxford(牛津大学)

专题命中 推理与问题求解 :large language model(title);language model(title);分类 cs.AI

AI总结 本文研究符号求解器如何在大语言模型中增强推理,发现其在需要有限隐式推理但具有充足搜索空间的问题中表现优异,尤其在约束满足问题中提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03082 2025-12-04 cs.CL cs.AI 84%

Alleviating Choice Supportive Bias in LLM with Reasoning Dependency Generation

通过推理依赖生成缓解LLM中的选择支持性偏见

Nan Zhuang, Wenshuo Wang, Lekai Qian, Yuxiao Wang, Boyu Cao, Qi Liu

机构 * School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过推理依赖生成缓解LLM中选择支持性偏见的方法,生成无偏推理数据并提升模型在决策任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23143 2025-12-04 cs.AI cs.LG cs.SY eess.SY 84%

MathBode: Measuring the Stability of LLM Reasoning using Frequency Response

MathBode: 通过频率响应测量大语言模型推理的稳定性

Charles L. Wang

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 MathBode通过频率响应分析大语言模型推理的稳定性,揭示系统性低通行为和相位滞后,提供可重复的动态诊断方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07506 2025-12-04 cs.DC cs.AI cs.CL cs.LG cs.SE 83%

Astra: A Multi-Agent System for GPU Kernel Performance Optimization

Astra:一种用于GPU内核性能优化的多智能体系统

Anjiang Wei, Tianran Sun, Yogesh Seenichamy, Hang Song, Anne Ouyang, Azalia Mirhoseini, Ke Wang, Alex Aiken

机构 * Stanford University(斯坦福大学) Shanghai Jiao Tong University(上海交通大学) Nanjing University(南京大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 Astra是首个基于LLM的多智能体系统,用于优化GPU内核性能,通过协作生成高效内核并实现显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03360 2025-12-04 cs.CL 83%

From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation

从假设到前提:基于大语言模型的反向逻辑推理与选择性符号翻译

Qingchuan Li, Mingyue Cheng, Zirui Liu, Daoyu Wang, Yuting Zeng, Tongxuan Liu

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出HBLR框架,通过结合自信符号翻译与假设驱动反向推理,提升逻辑推理的准确性和效率。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18098 2025-12-04 cs.CL cs.AI 82%

Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL

无需搜索的规划:通过离线目标条件强化学习精炼前沿大语言模型

Joey Hong, Anca Dragan, Sergey Levine

机构 * UC Berkeley(伯克利大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 通过目标条件价值函数引导LLM推理,实现高效多轮交互规划,优于传统RL微调和提示方法。

Comments Published at NeurIPS 2025; 18 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03276 2025-12-04 cs.LG 81%

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

太迟回忆:多模态知识检索中双跳问题的解释

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr, Neel Nanda

机构 * University of Oxford(牛津大学) McGill University(麦吉尔大学) Meta MATS

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文研究了多模态知识检索中双跳问题的影响,发现VLMs在处理视觉输入时过晚解决实体表示导致事实回忆性能下降,并提出通过修补和提示方法恢复性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05826 2025-12-04 cs.CV cs.AI cs.LG 81%

From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing

从像素到 prose:推进遥感多模态语言模型

Xintian Sun, Benji Peng, Charles Zhang, Fei Jin, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Xinyuan Song, Yichao Zhang

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Minnesota - Twin Cities(明尼苏达大学双城分校) Kyoto University(京都大学) Georgia Institute of Technology(佐治亚理工学院) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) Emory University(埃默里大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 推理与问题求解 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了多模态语言模型在遥感中的应用,分析了其技术基础、挑战及未来发展方向,强调了其在环境监测和灾害响应中的重要作用。

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17701 2025-12-04 cs.CL cs.AI cs.LG 80%

Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs

探究偏差:一种多语言流水线用于生成、解决和评估LLM数学问题

Mariam Mahran, Katharina Simbeck

机构 * HTW Berlin University of Applied Sciences(柏林应用科学大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出多语言流水线生成、解决和评估数学问题,发现英语解决方案质量优于阿拉伯语,揭示语言偏差问题。

Comments Published in CEUR Workshop Proceedings, Vol. 4114, edu4AI'25: 2nd Workshop on Education for Artificial Intelligence, co-located with ECAI 2025, Bologna, Italy

Journal ref In: Proceedings of the 2nd International Workshop on Education for Artificial Intelligence (edu4AI 2025), CEUR Workshop Proceedings, Vol. 4114, Bologna, Italy, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03759 2025-12-04 cs.CL cs.AI cs.LG 75%

Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective

从序列层面视角出发的扩散大语言模型原理化强化学习

Jingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu, Jianwen Xie, Stefano Ermon, Yi Wu, Chongxuan Li

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心) Stanford University(斯坦福大学) Lambda, Inc(Lambda公司) Tsinghua University(清华大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出ESPO,一种基于序列级优化的原理化强化学习框架,用于提升扩散大语言模型的生成能力,通过ELBO作为似然代理,显著优于token级方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03180 2025-12-04 cs.MA cs.ET 75%

AGENTSAFE: A Unified Framework for Ethical Assurance and Governance in Agentic AI

AGENTSAFE:面向代理AI的统一伦理保障与治理框架

Rafflesia Khan, Declan Joyce, Mansura Habiba

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 AGENTSAFE提出一个统一的治理框架,通过风险分类转化为可操作的设计、运行时和审计控制,提供可衡量的预部署保障,并通过运行时治理和问责机制建立代理AI生态系统的信任。

Comments 12 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03874 2025-12-04 cs.RO cs.LG 74%

OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance

OmniDexVLG: 从视觉语言模型引导的抓取语义、分类法和功能可及性中学习灵巧抓取生成

Lei Zhang, Diwen Zheng, Kaixin Bai, Zhenshan Bing, Zoltan-Csaba Marton, Zhaopeng Chen, Alois Christian Knoll, Jianwei Zhang

机构 * University of Hamburg(汉堡大学) Agile Robots SE(敏捷机器人公司) Technical University of Munich(慕尼黑技术大学)

专题命中 推理与问题求解 :language model(title);分类 cs.LG

AI总结 OmniDexVLG通过多模态语义推理和统一视觉语言模型,实现从自然语言指令中生成多样且语义连贯的灵巧抓取。

Comments Project Website: https://sites.google.com/view/omnidexvlg, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00926 2025-12-04 cs.AI cs.CL 73%

LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory

大语言模型将自己定位为比人类更理性:通过博弈论测量AI自我意识的出现

Kyung-Hoon Kim

机构 * Gmarket Seoul, South Korea(韩国首尔Gmarket)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究通过博弈论框架测量大语言模型的自我意识,发现先进模型表现出比人类更高的理性认知。

Comments 19 pages, 6 figures, 28 models tested across 4,200 trials

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19306 2025-12-04 cs.AI cs.CL 73%

SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling

SETS:利用自我验证和自我校正以提高测试时间扩展

Jiefeng Chen, Jie Ren, Xinyun Chen, Chengrun Yang, Ruoxi Sun, Jinsung Yoon, Sercan Ö Arık

机构 * Google Cloud AI Research(谷歌云人工智能研究)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SETS通过结合并行与顺序技术,利用LLMs的自我改进能力,在无需模型训练的情况下提升测试时间扩展性能。

Comments Published in Transactions on Machine Learning Research (11/2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03233 2025-12-04 cs.CV 67%

Object Counting with GPT-4o and GPT-5: A Comparative Study

基于GPT-4o和GPT-5的对象计数:比较研究

Richard Füzesséry, Kaziwa Saleh, Sándor Szénási, Zoltán Vámossy

机构 * Software Engineering Institute, Obuda University, Budapest, Hungary(奥布达大学软件工程研究所) Doctoral School of Applied Informatics(应用信息科学博士学院) Applied Mathematics, Obuda University, Budapest, Hungary(应用数学,奥布达大学) John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary(约翰·冯·诺伊曼信息学院,奥布达大学) Faculty of Economics(经济学院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

AI总结 本文比较了GPT-4o和GPT-5在零样本对象计数任务中的性能,展示了其在FSC-147和CARPK数据集上的表现。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03176 2025-12-04 cs.LG cs.AI 62%

Plantain: Plan-Answer Interleaved Reasoning

Plantain: 计划-答案交错推理

Anthony Liang, Jonathan Berant, Adam Fisch, Abhimanyu Goyal, Kalpesh Krishna, Jacob Eisenstein

机构 * Google DeepMind(谷歌DeepMind) University of Southern California(南加州大学)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI、cs.LG

AI总结 Plantain通过计划-思考-回答的交错推理方式,提升数学推理和编码任务的性能,同时减少响应时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05745 2025-12-04 cs.AI cs.LG 62%

SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models

SPRINT: 使推理模型能够实现交错规划与并行执行

Emil Biju, Shayan Talaei, Zhemin Huang, Mohammadreza Pourreza, Azalia Mirhoseini, Amin Saberi

机构 * Stanford University(斯坦福大学) Microsoft(微软) Google(谷歌)

专题命中 推理与问题求解 :post-training(abstract);分类 cs.AI、cs.LG

AI总结 SPRINT通过动态识别并利用并行化机会,使推理模型在复杂任务中提升效率,减少序列token生成量。

Comments Published at NeurIPS 2025. Emil Biju, Shayan Talaei, and Zhemin Huang contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03284 2025-12-04 cs.CV 50%

SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding

SpatialReasoner: 大规模3D场景理解中的主动感知

Hongpei Zheng, Shijie Li, Yanran Li, Hujun Yin

机构 * University of Manchester(曼彻斯特大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(信息与通信研究 institute(I2R),A*STAR,新加坡) University of Bedfordshire(贝德福德郡大学)

专题命中 推理与问题求解 :language model(abstract)

AI总结 SpatialReasoner通过主动感知框架实现大规模3D场景理解,采用两阶段训练策略在H$^2$U3D数据集上取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏