arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10577 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10577 篇

2405.16506 2025-07-15 cs.LG 57%

GRAG: Graph Retrieval-Augmented Generation

Yuntong Hu, Zhihan Lei, Zheng Zhang, Bo Pan, Chen Ling, Liang Zhao

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07735 2025-07-14 cs.LG 57%

Assessing the Chemical Intelligence of Large Language Models

Nicholas T. Runcie, Charlotte M. Deane, Fergus Imrie

机构 * Department of Statistics University of Oxford(统计系牛津大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06323 2025-07-10 cs.CR cs.AI 57%

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms

Tarek Gasmi, Ramzi Guesmi, Ines Belhadj, Jihene Bennaceur

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12464 2025-07-10 cs.CL 57%

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, Maarten Sap

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Published at NAACL 2025, Albuquerque, New Mexico, USA

Journal ref Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) 2373-2403

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06157 2025-07-09 cs.RO cs.CL 57%

Evaluation of Habitat Robotics using Large Language Models

William Li, Lei Hamilton, Kaise Al-natour, Sanjeev Mohindra

机构 * Artificial Intelligence Group MIT Lincoln Lab(人工智能组 MIT 林肯实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 6 pages, IEEE HPEC submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05997 2025-07-09 cs.CL 57%

DocIE@XLLM25: In-Context Learning for Information Extraction using Fully Synthetic Demonstrations

Nicholas Popovič, Ashish Kangen, Tim Schopf, Michael Färber

机构 * ScaDS.AI & TU Dresden(ScaDS.AI与德累斯顿技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05716 2025-07-09 cs.AI 57%

Divergent Realities: A Comparative Analysis of Human Expert vs. Artificial Intelligence Based Generation and Evaluation of Treatment Plans in Dermatology

Dipayan Sengupta, Saumya Panda

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 13 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04697 2025-07-08 cs.LG cs.DC cs.MS 57%

Performance Evaluation of General Purpose Large Language Models for Basic Linear Algebra Subprograms Code Generation

Daichi Mukunoki, Shun-ichiro Hayashi, Tetsuya Hoshino, Takahiro Katagiri

机构 * Information Technology Center\ University Nagoya, Aichi 466-8601 Email Graduate School of Informatics\ University Nagoya, Aichi 466-8601 Email

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 8 pages, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04333 2025-07-08 cs.CV cs.CL 57%

Computed Tomography Visual Question Answering with Cross-modal Feature Graphing

Yuanhe Tian, Chen Su, Junwen Duan, Yan Song

机构 * University of Washington(华盛顿大学) University of Science and Technology of China(中国科学技术大学) Central South University(中南大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03834 2025-07-08 cs.AI 57%

Economic Evaluation of LLMs

Michael J. Zellinger, Matt Thomson

机构 * California Institute of Technology(加利福尼亚技术研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14727 2025-07-08 cs.RO cs.AI 57%

Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models

Huihan Liu, Rutav Shah, Shuijing Liu, Jack Pittenger, Mingyo Seo, Yuchen Cui, Yonatan Bisk, Roberto Martín-Martín, Yuke Zhu

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03477 2025-07-08 cs.AI 57%

REAL: Benchmarking Abilities of Large Language Models for Housing Transactions and Services

Kexin Zhu, Yang Han

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03410 2025-07-08 cs.CL cs.DB cs.ET 57%

Graph Repairs with Large Language Models: An Empirical Study

Hrishikesh Terdalkar, Angela Bonifati, Andrea Mauri

机构 * Lyon1 University, CNRS LIRIS France Lyon1 University, CNRS LIRIS \& IUF France Lyon1 University, CNRS LIRIS Lyon1 University, CNRS LIRIS \& IUF

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted to the 8th GRADES-NDA 2025 @ SIGMOD/PODS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03013 2025-07-08 cs.CY cs.AI 57%

Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison

Aymeric de Chillaz, Anna Sotnikova, Patrick Jermann, Antoine Bosselut

机构 * EPFL(苏黎世联邦理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21053 2025-07-08 cs.CL 57%

MT2-CSD: A New Dataset and Multi-Semantic Knowledge Fusion Method for Conversational Stance Detection

Fuqiang Niu, Genan Dai, Yisha Lu, Jiayu Liao, Xiang Li, Hu Huang, Bowen Zhang

机构 * College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳技术大学) University of Washington(华盛顿大学) University of Science and Technology of China(中国科学技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02256 2025-07-04 cs.LG cs.RO 57%

Uncertainty-aware Reward Design Process

Yang Yang, Xiaolu Zhou, Bosong Ding, Miao Xin

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 34 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01627 2025-07-03 cs.CL 57%

Chart Question Answering from Real-World Analytical Narratives

Maeve Hutchinson, Radu Jianu, Aidan Slingsby, Jo Wood, Pranava Madhyastha

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments This paper has been accepted to the ACL Student Research Workshop (SRW) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01282 2025-07-03 cs.AI cs.HC 57%

Beyond Black-Box AI: Interpretable Hybrid Systems for Dementia Care

Matthew JY Kang, Wenli Yang, Monica R Roberts, Byeong Ho Kang, Charles B Malpas

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22405 2025-07-03 cs.CL 57%

Sequential Diagnosis with Language Models

Harsha Nori, Mayank Daswani, Christopher Kelly, Scott Lundberg, Marco Tulio Ribeiro, Marc Wilson, Xiaoxuan Liu, Viknesh Sounderajah, Jonathan Carlson, Matthew P Lungren, Bay Gross, Peter Hames, Mustafa Suleyman, Dominic King, Eric Horvitz

机构 * Microsoft AI(微软人工智能)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12084 2025-07-03 cs.CL 57%

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Jianshu Zhang, Dongyu Yao, Renjie Pi, Paul Pu Liang, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港科技大学) Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Project Page: https://vlm2-bench.github.io/ Camera Ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00718 2025-07-02 cs.CL 57%

AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation

Elizabeth Fons, Elena Kochkina, Rachneet Kaur, Zhen Zeng, Berowne Hlavaty, Charese Smiley, Svitlana Vyetrenko, Manuela Veloso

机构 * J.P. Morgan AI Research(摩根大通人工智能研究)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00525 2025-07-02 cs.CV cs.AI 57%

Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving

Djamahl Etchegaray, Yuxia Fu, Zi Huang, Yadan Luo

机构 * The University of Queensland(昆士兰大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14518 2025-07-02 eess.AS cs.CL cs.SD 57%

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Chun-Yi Kuan, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering(通信工程研究生院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted to Interspeech 2025. Project Website: https://kuan2jiu99.github.io/Balsa

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23725 2025-07-01 cs.RO cs.AI 57%

PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?

Atharva Gundawar, Som Sagar, Ransalu Senanayake

机构 * School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23122 2025-07-01 cs.CL cs.CY 57%

Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models

Shivam Sharma, Tanmoy Chakraborty

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20832 2025-07-01 cs.CV cs.AI 57%

Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models

Cansu Korkmaz, Ahmet Murat Tekalp, Zafer Dogan

机构 * Department of Electrical and Electronics Engineering and KUIS AI Center, Koç University(电气与电子工程系和KUIS人工智能中心,科克大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 14 pages, 9 figures, 5 tables, accepted to IEEE Transactions on Circuits and Systems for Video Technology

Journal ref IEEE Transactions on Circuits and Systems for Video Technology 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22026 2025-06-30 cs.IR cs.AI 57%

Literature-Grounded Novelty Assessment of Scientific Ideas

Simra Shahid, Marissa Radensky, Raymond Fok, Pao Siangliulue, Daniel S. Weld, Tom Hope

机构 * Microsoft(微软公司) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21784 2025-06-30 cs.AI 57%

MobiVerse: Scaling Urban Mobility Simulation with Hybrid Lightweight Domain-Specific Generator and Large Language Models

Yifan Liu, Xishun Liao, Haoxuan Ma, Jonathan Liu, Rohan Jadhav, Jiaqi Ma

机构 * UCLA Mobility Lab under the Department of Civil and Environmental Engineering, University of California, Los Angeles(加州大学洛杉矶分校土木与环境工程系移动实验室)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11095 2025-06-30 cs.CL 57%

A Survey of Large Language Models in Psychotherapy: Current Landscape and Future Directions

Hongbin Na, Yining Hua, Zimu Wang, Tao Shen, Beibei Yu, Lilin Wang, Wei Wang, John Torous, Ling Chen

机构 * Australian Artificial Intelligence Institute(澳大利亚人工智能研究所) University of Technology Sydney(悉尼科技大学) Harvard University(哈佛大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10573 2025-06-27 cs.CY cs.LG 57%

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Olawale Salaudeen, Anka Reuel, Ahmed Ahmed, Suhana Bedi, Zachary Robertson, Sudharsan Sundar, Ben Domingue, Angelina Wang, Sanmi Koyejo

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Correspondence to olawale@mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏