arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-11-20 至 2025-11-20 共收录 59 信号源:cs.CL, cs.AI, cs.LG

1. 复杂问题求解 9 篇

2511.12344 2025-11-20 cs.AI 79%

Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning

Baolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong, Lingrui Mei, Yuyao Ge, Yilong Xu, Jiafeng Guo, Xueqi Cheng

机构 * Key Laboratory of Network Data Science(网络数据科学实验室) Technology, Institute of Computing Technology, Chinese Academy of Sciences(技术,计算技术研究所,中国科学院) State Key Laboratory of AI Safety(人工智能安全国家重点实验室) University of Chinese Academy of Sciences(中国科学院大学) University of California, Merced(加州大学梅尔塞德斯分校)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16983 2025-11-20 cs.CL cs.AI 76%

ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation

Riccardo Pozzi, Matteo Palmonari, Andrea Coletta, Luigi Bellomarini, Jens Lehmann, Sahar Vahdati

机构 * University of Milano-Bicocca(米兰-比科卡大学) Banca d'Italia(意大利银行) Institute for Applied Informatics (InfAI)(应用信息研究所) TIB Leibniz Information Centre for Science and Technology(科学与技术信息中心) ScaDS.AI Dresden/Leipzig(ScaDS.AI 德累斯顿/莱比锡分校) Technische Universität Dresden(德累斯顿技术大学) Data Science Institute(数据科学研究所) Leibniz University Hannover(汉诺威莱布尼茨大学) Amazon(亚马逊)

专题命中 复杂问题求解 :reasoning(title);分类 cs.CL、cs.AI

Comments 19 pages, 6 figures, accepted at ISWC

Journal ref The Semantic Web - ISWC 2025. ISWC 2025. Lecture Notes in Computer Science, vol 16140. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15342 2025-11-20 cs.HC cs.AI 57%

Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions

Shan Shan

机构 * Zhejiang University(浙江大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02249 2025-11-20 cs.RO cs.AI 57%

Natural Selection via Foundation Models for Soft Robot Evolution

Changhe Chen, Xiaohao Xu, Xiangdong Wang, Xiaonan Huang

机构 * University of Michigan-Ann Arbor(密歇根大学安娜堡分校)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10596 2025-11-20 eess.SY cs.SY eess.SP 50%

Artifacts Are Not Noise: Embodied Resonance and the 70% Signal Loss in Conventional EEG

Ahmed Gamal Eldin

专题命中 复杂问题求解 :reasoning(abstract)

Comments v3: Major revision. Adds ICA vs. raw data comparison. Finds artifact rejection masks the core correlation (r=0.59 -> r=0.19), supporting an embodied cognition framework. R vs. ITC comparison added to main text. Expanded discussion on methodology, embodied resonance, & AI. Abstract & theory updated based on new data & feedback

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 推理评测 13 篇

2511.14777 2025-11-20 cs.AI 79%

The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs

Mahdi Samiei, Mahdi Mansouri, Mahdieh Soleymani Baghshah

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08008 2025-11-20 cs.AI 79%

Combining LLM Semantic Reasoning with GNN Structural Modeling for Multi-View Multi-Label Feature Selection

Zhiqi Chen, Yuzhou Liu, Jiarui Liu, Wanfu Gao

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10900 2025-11-20 cs.CL cs.AI 79%

Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering

Xueren Ge, Sahil Murtaza, Anthony Cortez, Homa Alemzadeh

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14214 2025-11-20 cs.AI cs.LG 62%

Do Large Language Models (LLMs) Understand Chronology?

Pattaraphon Kenny Wongchamcharoen, Paul Glasserman

机构 * University of California, Berkeley(加州大学伯克利分校) Columbia Business School(哥伦比亚大学商学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Version 2: corrected footnote and added code repository link. Extended version of our work presented at the AAAI-26 AI4TS Workshop (poster) and AAAI-26 Student Abstract Program (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15061 2025-11-20 cs.AI cs.IR cs.LG 62%

Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering

Haodong Chen, Guido Zuccon, Teerapong Leelanupab

机构 * The University of Queensland(昆士兰大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted to SIGIR-AP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15567 2025-11-20 cs.CV cs.CL cs.HC 57%

Computer-Use Agents as Judges for Generative User Interface

计算机使用代理作为生成用户界面的法官

Kevin Qinghong Lin, Siyuan Hu, Linjie Li, Zhengyuan Yang, Lijuan Wang, Philip Torr, Mike Zheng Shou

机构 * University of Oxford(牛津大学) Show Lab, National University of Singapore(新加坡国立大学Show实验室) Microsoft(微软公司)

专题命中 推理评测 :verifier(abstract);分类 cs.CL

AI总结 本研究提出Coder-CUA协作框架,通过代理作为法官与编码模型协作,提升自动GUI设计的效率和可靠性。

Comments Project: https://showlab.github.io/AUI Github: https://github.com/showlab/AUI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14903 2025-11-20 cs.LG cs.SE 57%

It's LIT! Reliability-Optimized LLMs with Inspectable Tools

Ruixin Zhang, Jon Donnelly, Zhicheng Guo, Ghazal Khalighinejad, Haiyang Huang, Alina Jade Barnett, Cynthia Rudin

机构 * Department of Computer Science(计算机科学系) Duke University(杜克大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Multi-Turn Interactions in Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14780 2025-11-20 cs.AI 57%

Ask WhAI:Probing Belief Formation in Role-Primed LLM Agents

Keith Moore, Jun W. Kim, David Lyu, Jeffrey Heo, Ehsan Adeli

机构 * Department of Biomedical Data Science, Stanford University(生物医学数据科学系,斯坦福大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Preprint. Accepted for publication at AIAS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14439 2025-11-20 cs.CL 57%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13889 2025-11-20 cs.CV cs.LG 57%

Uni-Hema: Unified Model for Digital Hematopathology

Abdul Rehman, Iqra Rasool, Ayisha Imran, Mohsen Ali, Waqas Sultani

机构 * Information Technology University of Punjab(旁遮普信息科技大学) Chughtai Lab(楚格塔实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24709 2025-11-20 cs.CV 50%

IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?

Yang Chen, Minghao Liu, Yufan Shen, Yunwen Li, Tianyuan Huang, Xinyu Fang, Tianyu Zheng, Wenxuan Huang, Cheng Yang, Daocheng Fu, Jianbiao Mei, Rong Wu, Yunfei Zhao, Licheng Wen, Xuemeng Yang, Song Mao, Qunshu Lin, Zhi Yu, Yongliang Shen, Yu Qiao, Botian Shi

机构 * IWR-Bench Team(IWR-Bench团队)

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14937 2025-11-20 cs.CR 50%

CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs

Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan, Arman Zharmagambetov, Manzil Zaheer, Saeed Mahloujifar, Kamalika Chaudhuri

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00261 2025-11-20 cs.CV cs.HC 50%

Spot The Ball: A Benchmark for Visual Social Inference

Neha Balamurugan, Sarah Wu, Adam Chun, Gabe Gaw, Cristobal Eyzaguirre, Tobias Gerstenberg

机构 * Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他推理 11 篇

2511.12001 2025-11-20 cs.CL cs.HC 89%

Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations

Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, Maarten Sap, Motahhare Eslami

机构 * Seoul National University(首尔国立大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

Comments Under review; 16 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20098 2025-11-20 cs.CL cs.AI 81%

Leveraging the Power of Large Language Models in Entity Linking via Adaptive Routing and Targeted Reasoning

Yajie Li, Albert Galimov, Mitra Datta Ganapaneni, Pujitha Thejaswi, De Meng, Priyanshu Kumar, Saloni Potdar

机构 * College of Information and Computer Sciences, University of Massachusetts Amherst(信息与计算机科学学院,马萨诸塞大学阿默斯特分校) Apple(苹果公司)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15137 2025-11-20 cs.LG cs.AI 81%

From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs

Xiaoxuan Wang, Bo Liu, Song Jiang, Jingzhou Liu, Jingyuan Qi, Xia Chen, Baosheng He

机构 * Meta

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15375 2025-11-20 cs.LG cs.AI 62%

Parameter Importance-Driven Continual Learning for Foundation Models

Lingxiang Wang, Hainan Zhang, Zhiming Zheng

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14010 2025-11-20 cs.CL cs.AI 62%

Knowledge-Grounded Agentic Large Language Models for Multi-Hazard Understanding from Reconnaissance Reports

Chenchen Kuai, Zihao Li, Braden Rosen, Stephanie Paal, Navid Jafari, Jean-Louis Briaud, Yunlong Zhang, Youssef M. A. Hashash, Yang Zhou

机构 * organization= Department One , addressline= Address One , city= City One , postcode= 00000 , state= State One , country= Country One organization= Department Two , addressline= Address Two , city= City Two , postcode= 22222 , state= State Two , country= Country Two organization= Zachry Department of Civil \& Environmental Engineering, Texas A\&M University , addressline= 3136 TAMU , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Engineering Technology Industrial Distribution, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Civil Environmental Engineering, University of Illinois Urbana-Champaign , city= Urbana , postcode= 61801 , state= IL , country= USA

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14772 2025-11-20 cs.CL cs.AI 62%

Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective

Zhuoyi Yang, Xu Guo, Tong Zhang, Huijuan Xu, Boyang Li

机构 * Department of Computer Science and Engineering, The Pennsylvania State University(宾夕法尼亚州立大学计算机科学与工程系) School of Computer Science and Engineering, Nanyang Technological University(南洋理工大学计算机科学与工程学院)

专题命中 其他推理 :chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08146 2025-11-20 cs.CL cs.LG 62%

Bias after Prompting: Persistent Discrimination in Large Language Models

Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi, Krishna Patel, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff

机构 * Apple(苹果公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15378 2025-11-20 cs.AI 57%

Terra Nova: A Comprehensive Challenge Environment for Intelligent Agents

Trevor McInroe

机构 * The University of Edinburgh(爱丁堡大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15389 2025-11-20 cs.IR 50%

Unveiling Inference Scaling for Difference-Aware User Modeling in LLM Personalization

Suyu Chen, Yimeng Bai, Yulong Huang, Xiaoyan Zhao, Yang Zhang

专题命中 其他推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15206 2025-11-20 cs.CR cs.IT math.IT 50%

Trustworthy GenAI over 6G: Integrated Applications and Security Frameworks

Bui Duc Son, Trinh Van Chien, Dong In Kim

专题命中 其他推理 :reasoning(abstract)

Comments 8 pages, 5 figures. Submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14762 2025-11-20 cs.DB 50%

Castle: Causal Cascade Updates in Relational Databases with Large Language Models

Yongye Su, Yucheng Zhang, Zeru Shi, Bruno Ribeiro, Elisa Bertino

专题命中 其他推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏