arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 18857 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 18857 篇

2510.15987 2026-02-17 cs.LG cs.AI 86%

Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models

语言模型中推理的算法原语与组合几何

Samuel Lippl, Thomas McGee, Kimberly Lopez, Ziwen Pan, Pierce Zhang, Salma Ziadi, Oliver Eberle, Ida Momennejad

机构 * Microsoft Research NYC(微软研究院纽约分部) Center for Theoretical Neuroscience(理论神经科学中心) Department of Psychology(心理学系) University of California Los Angeles(加州大学洛杉矶分校) Institute for Pure and Applied Mathematics(纯粹与应用数学研究所) Emory University(埃默里大学) Rice University(里士满大学) Mount Holyoke College(马里兰霍克学院) Technische Universität Berlin(柏林技术大学) BIFOLD-Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)

专题命中 推理与问题求解 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 研究通过分析语言模型中的算法原语,揭示其推理过程的组合几何结构,并展示原语在跨任务和跨模型中的可迁移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12285 2026-02-16 cs.CL cs.AI 86%

From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness

从有偏聊天机器人到有偏代理:检验角色分配对LLM代理鲁棒性的影响

Linbo Cao, Lihao Sun, Yang Yue

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究发现基于人口统计学的人设分配会影响LLM代理的行为和性能,导致高达26.2%的降级,揭示了代理系统中隐含偏见和行为波动的风险。

Comments Accepted to the AAAI 2026 TrustAgent Workshop. 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07695 2026-02-12 cs.AI cs.CL cs.IR cs.MM 86%

EventCast: Hybrid Demand Forecasting in E-Commerce with LLM-Based Event Knowledge

EventCast:基于LLM事件知识的电子商务混合需求预测

Congcong Hu, Yuang Shi, Fan Huang, Yang Xiang, Zhou Ye, Ming Jin, Shiyu Wang

机构 * ByteDance China(字节跳动中国) National University of Singapore(新加坡国立大学) Griffith University(格里菲斯大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 EventCast通过整合LLM事件知识提升电子商务需求预测精度,实现86.9%和97.7%的MAE和MSE改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04326 2026-02-05 cs.AI cs.CL cs.MA 86%

From Assumptions to Actions: Turning LLM Reasoning into Uncertainty-Aware Planning for Embodied Agents

从假设到行动:将大语言模型推理转化为具有不确定性的代理规划

SeungWon Seo, SooBin Lim, SeongRae Noh, Haneul Kim, HyeongYeop Kang

机构 * Department of Computer Science and Engineering, Korea University(计算机科学与工程系,韩国大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 PCE框架通过结构化决策树将LLM推理中的假设转化为可靠策略,提升多智能体环境下的规划效率和任务完成率。

Comments 31 pages, 10 figures, Accepted ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03608 2026-02-04 cs.CL cs.AI cs.IR 86%

Controlling Output Rankings in Generative Engines for LLM-based Search

在基于大语言模型的搜索生成引擎中控制输出排名

Haibo Jin, Ruoxi Chen, Peiyan Zhang, Yifeng Luo, Huimin Zeng, Man Luo, Haohan Wang

机构 * School of Information Sciences, University of Illinois at Urbana-Champaign, IL, USA(伊利诺伊大学厄巴纳-香槟分校信息科学学院) Independent Researcher, Starc Institute(Starc研究所独立研究者) Research Scientist, Intel Labs, Santa Clara, CA, USA(英特尔实验室研究科学家)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 CORE通过优化内容生成,提升基于LLM的搜索中小企业产品的可见性,实现高推广成功率。

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17813 2026-02-04 cs.CL cs.AI 86%

Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning

不要过度思考。为改进大语言模型推理而偏好更短的思考链

Michael Hassid, Gabriel Synnaeve, Yossi Adi, Roy Schwartz

机构 * FAIR Team, Meta(Meta FAIR团队) The Hebrew University of Jerusalem(耶路撒冷希伯来大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出short-m@k方法,通过并行生成较短的思考链提高大语言模型推理效率,实验表明其在低计算环境下表现优于传统多数投票。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02386 2026-02-03 cs.AI cs.IR cs.LG 86%

Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing

信任由设计:基于技能配置的透明、成本感知的LLM路由

Mika Okamoto, Ansel Kaplan Erol, Glenn Matlin

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 BELLA框架通过基于技能的可解释方法,为LLM路由提供透明、成本感知的最优模型选择方案,帮助从业者在性能与成本间做出合理权衡。

Comments Appeared at MLSys YPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01034 2026-02-03 cs.AI cs.CL 86%

Discovering Process-Outcome Credit in Multi-Step LLM Reasoning

在多步骤LLM推理中发现过程-结果信用

Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Saman Halgamuge

机构 * The University of Melbourne(墨尔本大学)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);SFT(abstract)

AI总结 本文提出了一种新的框架,通过分步边际信息增益机制和解耦掩码策略,提升多步骤LLM推理的样本效率和准确性,并增强模型的分布外鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05266 2026-02-03 cs.AR cs.CL cs.LG 86%

Understanding and Mitigating Errors of LLM-Generated RTL Code

理解并缓解LLM生成的RTL代码错误

Jiazheng Zhang, Cheng Liu, Long Cheng, Xiaowei Li, Huawei Li

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出基于LLM的框架,通过检索增强生成、规则检查、多模态转换和迭代仿真调试,显著提升了RTL代码生成的准确性。

Comments Accepted by IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00016 2026-02-03 cs.CL cs.AI 86%

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

PTCBENCH: 对LLM系统中人格特质情境稳定性的基准测试

Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang, Junjie Wang, Qiang Hu, Chao Shen, Xiaofei Xie

机构 * Singapore Management University(新加坡管理大学) Tianjin University(天津大学) Nanyang Technological University(南洋理工大学) Xi’an Jiaotong University(西安交通大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 PTCBENCH通过评估LLM在不同情境下的人格一致性,揭示外部事件对LLM人格和推理能力的影响,为构建更稳健的心理学对齐AI系统提供新视角。

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18292 2026-02-02 cs.LG cs.AI 86%

TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment

TriPlay-RL:三角色自我对抗强化学习用于大语言模型安全对齐

Zhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Duohe Ma, Xiangzheng Zhang, Tong Yang, Lin Sun

机构 * Peking University(北京大学) Qiyuan Tech(启元科技) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 TriPlay-RL通过三角色自我对抗强化学习,实现LLM安全对齐的高效且可扩展范式,提升攻击有效性、安全性能和评估准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11358 2026-01-28 cs.CL cs.AI cs.IR 86%

LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation

针对大语言模型的特定效用:检索增强生成的新视角

Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Baidu Inc(百度公司)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出LLM特定效用的概念,指出不同大语言模型对证据的需求不同,提出构建基准以研究这种效用,并推动生成器定制的证据选择方法改进RAG。

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10132 2026-01-27 cs.AI cs.LG 86%

Is More Context Always Better? Examining LLM Reasoning Capability for Time Interval Prediction

更多上下文总是更好吗?检验LLM在时间间隔预测中的推理能力

Yanan Cao, Farnaz Fallahi, Murali Mohana Krishna Dandu, Lalitesh Morishetti, Kai Zhao, Luyi Ma, Sinduja Subramaniam, Jianpeng Xu, Evren Korpeoglu, Kaushiki Nag, Sushant Kumar, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文研究LLM在时间间隔预测中的推理能力,发现尽管LLM在某些任务中表现良好,但其在捕捉定量时间结构方面存在局限,且过多上下文反而可能降低性能。

Comments Accepted at The Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14456 2026-01-22 cs.AI cs.LG 86%

On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL

在LLM规划中的泛化差距:测试与验证者奖励强化学习

Valerio Belcamino, Nicholas Attolino, Alessio Capitanelli, Fulvio Mastrogiovanni

机构 * Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genoa(信息学、生物工程、机器人学和系统工程系,热那亚大学) AIKO S.r.l.(AIKO公司)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究发现微调LLM在规划任务中存在显著的泛化差距,通过三种诊断干预揭示模型依赖领域特定模式而非可转移能力。

Comments 9 pages, 4 figures, 3 tables, 2 pages of supplementary materials. Submitted to a conference implementing a double-blind review process

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13392 2026-01-21 cs.CL cs.AI cs.FL 86%

Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks

超越记忆:在未见的计算理论任务上测试LLM推理能力

Shlok Shelat, Jay Raval, Souvik Roy, Manas Gaur

机构 * Ahmedabad University Gujarat, India(古吉拉特邦阿赫迈德亚布大学) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文通过构建DFA基准测试,揭示LLM在未见计算理论任务上的推理缺陷,发现其在处理复杂约束和语义一致性时存在系统性不足。

Comments 30 pages, 11 figures, 6 tables, Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11905 2026-01-21 cs.AI cs.LG math.ST stat.TH 86%

LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning

LIBRA:基于语言模型的带状 recourse 算法用于个性化治疗计划

Junyu Cao, Ruijiang Gao, Esmaeil Keyvanshokooh, Jianhao Ma

机构 * McCombs School of Business, University of Texas at Austin(德克萨斯大学奥斯汀分校麦克斯韦商学院) Naveen Jindal School of Management, University of Texas at Dallas(德克萨斯大学达拉斯分校奈文·金达管理学院) Mays Business School, Texas A&M University(德克萨斯农工大学梅斯商学院) Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院)

专题命中 推理与问题求解 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 LIBRA 是一种结合大语言模型和带状学习的算法,用于在个性化治疗中实现更高效的决策和鲁棒性。

Comments 50 pages. Previous version with human-AI collaboration: arXiv:2410.14640

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09980 2026-01-16 cond-mat.mtrl-sci cs.AI cs.LG 86%

Performance of AI agents based on reasoning language models on ALD process optimization tasks

基于推理语言模型的AI代理在ALD过程优化任务中的性能

Angel Yanguas-Gil

机构 * Applied Materials Division, Argonne National Laboratory, Lemont, IL 60439, USA(阿贡国家实验室应用材料部)

专题命中 推理与问题求解 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 研究基于推理语言模型的AI代理在ALD过程优化中的性能,发现其能有效完成优化任务,但存在响应不确定性及潜在的推理偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17627 2026-01-14 cs.CL cs.AI 86%

The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis

思维的演变:通过推理动态分析追踪LLM过度思考

Zihao Wei, Liang Pang, Jiahao Liu, Wenjie Shi, Jingcheng Deng, Shicheng Xu, Zenghao Duan, Fei Sun, Huawei Shen, Xueqi Cheng

机构 * State Key Laboratory of AI Safety,Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 通过分析推理动态,提出RCPD方法以识别LLM推理完成点,减少token使用量同时保持准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04254 2026-01-09 cs.AI cs.LG 86%

Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models

中等规模语言模型中多跳上下文推理的扩展趋势

Brady Steele, Micah Katz

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 推理与问题求解 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 研究通过对比不同模型的多跳推理能力,发现多代理系统在推理任务中表现优于规则方法,且模型基础能力影响放大效果。

Comments 18 pages, 6 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19361 2026-01-09 cs.CL cs.AI 86%

AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation

AgenticMath: 通过基于代理的数学数据生成增强LLM推理

Xianyang Liu, Yilin Liu, Shuai Wang, Hao Cheng, Andrew Estornell, Yuzhi Zhao, Jun Shu, Jiaheng Wei

机构 * King’s College London(伦敦国王学院) Hong Kong Baptist University(香港 Baptist 大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ByteDance Seed(字节跳动种子) City University of Hong Kong(香港城市大学) Xi’an Jiaotong University(西安交通大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 AgenticMath通过基于代理的数学数据生成方法,提升LLM在数学推理任务上的性能。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23167 2025-12-30 cs.AI cs.LG cs.MA 86%

SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search

SPIRAL:通过 grounded 和 reflective 搜索实现符号 LLM 规划

Yifan Zhang, Giridhar Ganapavarapu, Srideepika Jayaraman, Bhavna Agrawal, Dhaval Patel, Achille Fokoue

机构 * IBM T.J. Watson Research Center(IBM TJ沃森研究中心)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 SPIRAL 通过 grounded 和 reflective 搜索实现更稳健高效的 LLM 规划

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22258 2025-12-30 cs.AI cs.LG cs.LO cs.SC 86%

Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method

逻辑草图提示(LSP):一种确定性和可解释的提示方法

Satvik Tripathi

专题命中 推理与问题求解 :prompting(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 逻辑草图提示(LSP)通过引入类型变量、确定性条件评估器和规则验证器,提升了大语言模型在需要严格规则遵守、确定性和可审计性任务中的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00617 2025-12-25 cs.CL cs.AI 86%

ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization

ART:自适应响应调节框架——基于多智能体竞赛的LLM响应优化方法

Omer Jauhar Khan

机构 * Department of Computer Science(计算机科学系) National University of Computer and Emerging Sciences (FAST-NUCES)(计算机与新兴科学国立大学(FAST-NUCES))

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 ART通过多智能体竞赛机制优化LLM响应,提升准确性和一致性,实现8.4%的质量提升和R²超过0.96的收敛效果。

Comments 14 pages, 11 figures, 5 tables. IEEE conference-style paper with appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04826 2025-12-24 cs.CL cs.AI 86%

Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History

大语言模型人格测量中的持久不稳定性:规模、推理和对话历史的影响

Tommaso Tosato, Saskia Helbling, Yorguin-Jose Mantilla-Ramos, Mahmood Hegazy, Alberto Tosato, David John Lemay, Irina Rish, Guillaume Dumas

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究发现大语言模型在人格测量中存在持续不稳定性,规模扩大、推理和对话历史等干预措施未能有效提升稳定性。

Comments Accepted at AAAI 2026, Track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16676 2025-12-19 cs.LG cs.CL 86%

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

DataFlow:面向数据导向AI时代的统一数据准备与工作流自动化框架

Hao Liang, Xiaochen Ma, Zhou Liu, Zhen Hao Wong, Zhengyang Zhao, Zimo Meng, Runming He, Chengyu Shen, Qifeng Cai, Zhaoyang Han, Meiyi Qiang, Yalin Feng, Tianyi Bai, Zewei Pan, Ziyi Guo, Yizhen Jiang, Jingwen Deng, Qijie You, Peichao Lai, Tianyu Guo, Chi Hsu Tsai, Hengyi Feng, Rui Hu, Wenkai Yu, Junbo Niu, Bohan Zeng, Ruichuan An, Lu Ma, Jihao Huang, Yaowei Zheng, Conghui He, Linpeng Tang, Bin Cui, Weinan E, Wentao Zhang

机构 * Peking University(北京大学) Institute for Advanced Algorithms Research(先进算法研究所) OriginHub Technology(OriginHub技术) OpenDataLab(OpenData实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) LLaMA-Factory Team(LLaMA-Factory团队)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 DataFlow通过统一的数据准备框架和自动化工作流,提升LLM性能,实现高效、可靠的数据处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15274 2025-12-18 cs.CL cs.AI 86%

Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning

开篇即胜,半途而废:基于前缀优化的强化学习用于大语言模型推理

Yiliu Sun, Zicheng Zhao, Yang Wei, Yanfang Zhang, Chen Gong

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 PPPO通过优化LLM推理的前缀部分,提升推理能力,实验显示其在推理任务中表现更优。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07785 2025-12-10 physics.data-an cs.AI cs.LG hep-ex 86%

Automating High Energy Physics Data Analysis with LLM-Powered Agents

利用LLM代理自动化高能物理数据分析

Eli Gendreau-Distler, Joshua Ho, Dongwon Kim, Luc Tomas Le Pottier, Haichen Wang, Chengxi Yang

机构 * Department of Physics, University of California, Berkeley, Berkeley, CA 94720, USA(加州大学伯克利分校物理系) Physics Division, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA(伯克利国家实验室物理部)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究利用LLM代理自动化高能物理数据分析,通过混合系统结合LLM和Snakemake工作流管理器,评估代理在多阶段工作流中的性能。

Comments 16 pages, 6 figures, 2 tables, the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) - Machine Learning and the Physical Sciences (ML4PS) workshop (poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20993 2025-12-09 cs.LG cs.AI 86%

Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning

子目标图增强的规划用于LLM引导的开放世界强化学习

Shanwei Fan, Bin Zhang, Zhiwei Xu, Yingxuan Teng, Siqi Dai, Lin Cheng, Guoliang Fan

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Shandong University(山东大学人工智能学院)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出SGA-ACR框架,通过整合环境特定的子目标图和多LLM规划流程,解决LLM在开放世界强化学习中的规划-执行对齐问题,提升子目标的可行性和可验证性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05998 2025-12-09 cs.AI cs.GT cs.LG 86%

Going All-In on LLM Accuracy: Fake Prediction Markets, Real Confidence Signals

全面投入LLM准确性:假预测市场,真实信心信号

Michael Todasco

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 通过将评估任务设计为赌博游戏,研究发现LLM的置信度信号可通过虚拟货币激励提升,但准确性提升未达统计显著性。

Comments 25 pages, 8 tables, 2 figures. Pilot study. Data, prompts, and code available at https://osf.io/dc24t/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18449 2025-12-02 cs.SE cs.AI cs.CL 86%

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

SWE-RL:通过强化学习提升大语言模型推理能力

Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, Sida I. Wang

机构 * Meta AI University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SWE-RL通过强化学习提升大语言模型在软件工程领域的推理能力,实现了41.0%的解决率,超越了现有中等规模LLM的性能。

Comments Accepted to NeurIPS 2025 Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏