arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10451 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10451 篇

2409.19492 2025-11-25 cs.CL cs.AI 62%

MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models

MedHalu:大型语言模型在医疗查询中生成幻觉的研究

Vibhor Agarwal, Yiqiao Jin, Mohit Chandra, Munmun De Choudhury, Srijan Kumar, Nishanth Sastry

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 MedHalu研究了LLM在医疗查询中生成幻觉的问题,提出MedHalu基准和MedHaluDetect框架,发现LLM在检测医学幻觉方面表现不佳,提出专家在环方法提升检测效果。

Comments Accepted at ICWSM2026. https://netsys.surrey.ac.uk/datasets/medhalu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17238 2025-11-24 cs.CL cs.AI cs.CV 62%

Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables

迷失在翻译与噪声中:对VLMs在现实表格上的失败模式的深入探讨

Anshul Singh, Rohan Chaudhary, Gagneet Singh, Abhay Kumary

机构 * Indian Institute of Science(印度科学研究院) Panjab University(旁遮普大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 MirageTVQA通过多语言和视觉不完美的表格挑战VLMs,揭示了模型在现实噪声和语言转换中的失败模式。

Comments Accepted as Spotligh Talk at EurIPS 2025 Workshop on AI For Tabular Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17309 2025-11-24 cs.AI cs.CL 62%

RubiSCoT: A Framework for AI-Supported Academic Assessment

RubiSCoT:一种支持人工智能的学术评估框架

Thorsten Fröhlich, Tim Schlippe

机构 * IU International University of Applied Sciences(国际应用科学大学)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 RubiSCoT通过人工智能技术提升学术论文评估的效率和一致性,提供从提案到最终提交的全面评估解决方案。

Journal ref The 6th International Conference on Artificial Intelligence in Education Technology (AIET 2025), Munich, Germany, 29-31 July 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16267 2025-11-21 cs.CL cs.AI 62%

From Confidence to Collapse in LLM Factual Robustness

从信心到崩溃:大语言模型事实鲁棒性的研究

Alina Fastowski, Bardh Prenkaj, Gjergji Kasneci

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出了一种通过分析令牌分布熵和温度缩放敏感性来评估大语言模型事实鲁棒性的新方法,并通过实验验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14214 2025-11-20 cs.AI cs.LG 62%

Do Large Language Models (LLMs) Understand Chronology?

Pattaraphon Kenny Wongchamcharoen, Paul Glasserman

机构 * University of California, Berkeley(加州大学伯克利分校) Columbia Business School(哥伦比亚大学商学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Version 2: corrected footnote and added code repository link. Extended version of our work presented at the AAAI-26 AI4TS Workshop (poster) and AAAI-26 Student Abstract Program (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15061 2025-11-20 cs.AI cs.IR cs.LG 62%

Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering

Haodong Chen, Guido Zuccon, Teerapong Leelanupab

机构 * The University of Queensland(昆士兰大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted to SIGIR-AP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14761 2025-11-19 cs.CV cs.AI cs.LG 62%

ARC Is a Vision Problem!

Keya Hu, Ali Cy, Linlu Qiu, Xiaoman Delores Ding, Runqian Wang, Yeyin Eva Zhu, Jacob Andreas, Kaiming He

机构 * MIT(麻省理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Technical Report. Project webpage: https://github.com/lillian039/VARC

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14631 2025-11-19 cs.CL cs.AI cs.CV cs.MA 62%

Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities

Kahaan Gandhi, Boris Bolliet, Inigo Zubeldia

机构 * Department of Physics, University of Cambridge, Cambridge, United Kingdom(剑桥大学物理系) Kavli Institute for Cosmology, University of Cambridge, Cambridge, United Kingdom(剑桥大学卡弗利天文研究所) Department of Physics and Astronomy, Haverford College, 370 Lancaster Avenue, Haverford, PA 19041, USA(哈弗福德学院物理与天文学系) Division of Physics, Mathematics and Astronomy, California Institute of Technology, Pasadena, CA 91125, USA(加州理工学院物理、数学与天文学系) Institute of Astronomy, University of Cambridge, Cambridge, United Kingdom(剑桥大学天文研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11654 2025-11-19 cs.IR cs.AI cs.CL 62%

FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection

Daniel Berhane Araya, Duoduo Liao

机构 * College of Engineering and Computing(工程与计算学院) George Mason University(乔治·马歇尔大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12443 2025-11-19 cs.SE cs.AI cs.DC cs.LG 62%

From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow

Sparsh Gupta, Kamalavasan Kamalakkannan, Maxim Moraru, Galen Shipman, Patrick Diehl

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) Franklin W. Olin College of Engineering(弗兰克林·W·奥林工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12851 2025-11-18 cs.CL cs.AI 62%

NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation

Kang Yin, Hye-Bin Shin

机构 * Dept. of Artificial Intelligence Korea University Seoul, Republic of Korea(人工智能系韩国大学首尔共和国韩国)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12769 2025-11-18 cs.AI cs.LG 62%

Event-CausNet: Unlocking Causal Knowledge from Text with Large Language Models for Reliable Spatio-Temporal Forecasting

Luyao Niu, Zepu Wang, Shuyi Guan, Yang Liu, Peng Sun

机构 * Duke Kunshan University, China(杜克大学昆山分校) Peking University, China(北京大学) Tongji University, China(同济大学) University of Ottawa, Canada(渥太华大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12630 2025-11-18 cs.CL cs.AI 62%

Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing

Maoqi Liu, Quan Fang, Yang Yang, Can Zhao, Kaiquan Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beihang University(北航) State Key Laboratory of CNS/ATM(国家空管流量管理技术实验室) Aviation Data Communication Corporation(航空数据通信公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to Advanced Engineering Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12472 2025-11-18 cs.CL cs.AI 62%

Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing

Mengying Wang, Chenhui Ma, Ao Jiao, Tuo Liang, Pengjun Lu, Shrinidhi Hegde, Yu Yin, Evren Gurkan-Cavusoglu, Yinghui Wu

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments The 40th AAAI Conference on Artificial Intelligence (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12460 2025-11-18 cs.LG cs.AI 62%

Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression Detection

Changzeng Fu, Shiwen Zhao, Yunze Zhang, Zhongquan Jian, Shiqi Zhao, Chaoran Liu

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments AAAI 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00091 2025-11-18 cs.AI cs.CL 62%

Ensemble Debates with Local Large Language Models for AI Alignment

Ephraiem Sarabamoun

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments The manuscript is being withdrawn to incorporate additional revisions and improvements

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12116 2025-11-18 cs.CL cs.AI 62%

LLMLagBench: Identifying Temporal Training Boundaries in Large Language Models

Piotr Pęzik, Konrad Kaczyński, Maria Szymańska, Filip Żarnecki, Zuzanna Deckert, Jakub Kwiatkowski, Wojciech Janowski

机构 * University of Lodz(洛兹大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11831 2025-11-18 cs.AI cs.CV cs.LG 62%

TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models

Wenhao Zhou, Hao Zheng, Rong Zhao

机构 * Center for Brain-Inspired Computing Research (CBICR)(脑启发计算研究中心) Department of Precision Instruments(精密仪器系) IDG/McGovern Institute for Brain Research(IDG/麦戈文脑研究学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08824 2025-11-18 cs.RO cs.AI cs.CL cs.CY 62%

LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions

Andrew Hundt, Rumaisa Azeem, Masoumeh Mansouri, Martim Brandão

机构 * Carnegie Mellon University(卡内基梅隆大学) King’s College London(伦敦国王学院) University of Birmingham(伯明翰大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Published in International Journal of Social Robotics (2025). 49 pages (65 with references and appendix), 27 Figures, 8 Tables. Andrew Hundt and Rumaisa Azeem are equal contribution co-first authors. The positions of the two co-first authors were swapped from arxiv version 1 with the written consent of all four authors. The Version of Record is available via DOI: 10.1007/s12369-025-01301-x

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15901 2025-11-17 cs.CL cs.AI 62%

Re-FRAME the Meeting Summarization SCOPE: Fact-Based Summarization and Personalization via Questions

Frederic Kirstein, Sonu Kumar, Terry Ruas, Bela Gipp

机构 * University of Göttingen(哥廷根大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10912 2025-11-17 cs.CL cs.AI 62%

Evaluating Large Language Models on Rare Disease Diagnosis: A Case Study using House M.D

Arsh Gupta, Ajay Narayanan Sridhar, Bonam Mingole, Amulya Yadav

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10881 2025-11-17 cs.CL cs.AI 62%

A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge

Jongyoon Song, Sangwon Yu, Sungroh Yoon

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments Accepted to IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10690 2025-11-17 cs.CL cs.AI 62%

Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games

Juntu Zhao, Jialing Zhang, Chongxuan Li, Dequan Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Renmin University of China(中国人民大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025 MTI-LLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19319 2025-11-14 cs.CL cs.AI 62%

FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering

Gyubok Lee, Elea Bach, Eric Yang, Tom Pollard, Alistair Johnson, Edward Choi, Yugang jia, Jong Ha Lee

机构 * Korea Advanced Institute of Science & Technology(韩国科学技术院) Verily Life Sciences(Verily 生物科技) Massachusetts Institute of Technology(麻省理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments ML4H 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09438 2025-11-13 cs.LG cs.AI 62%

LLM-Guided Dynamic-UMAP for Personalized Federated Graph Learning

Sai Puppala, Ismail Hossain, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder

机构 * University of Texas at El Paso(德克萨斯理工大学) Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09067 2025-11-13 cs.CL cs.AI 62%

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

Gailun Zeng, Ziyang Luo, Hongzhan Lin, Yuchen Tian, Kaixin Li, Ziyang Gong, Jianxiong Guo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) National University of Singapore(新加坡国立大学) Beijing Normal University(北京师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 14 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08242 2025-11-12 cs.AI cs.CL 62%

Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents

Waseem AlShikh, Muayad Sayed Ali, Brian Kennedy, Dmytro Mozolevskyi

机构 * Writer, Inc.(Writer公司)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07982 2025-11-12 cs.CL cs.AI 62%

NOTAM-Evolve: A Knowledge-Guided Self-Evolving Optimization Framework with LLMs for NOTAM Interpretation

Maoqi Liu, Quan Fang, Yuhao Wu, Can Zhao, Yang Yang, Kaiquan Cai

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14359 2025-11-12 cs.CV cs.AI cs.CL 62%

A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling

Kyle Buettner, Jacob T. Emmerson, Adriana Kovashka

机构 * Intelligent Systems Program(智能系统项目) Department of Computer Science(计算机科学系)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted at IJCNLP-AACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24409 2025-11-11 cs.CL cs.AI 62%

When Language Shapes Thought: Cross-Lingual Transfer of Factual Knowledge in Question Answering

Eojin Kang, Juae Kim

机构 * Hankuk University of Foreign Studies(韩国外国语大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted at CIKM2025 (Expanded version)

详情

展开后加载摘要…

URL PDF HTML 收藏