arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10549 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10549 篇

2510.26495 2025-11-14 cs.DB cs.CL 57%

Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration

Linzhuang Sun, Tianyu Guo, Hao Liang, Yuying Li, Qifeng Cai, Jingxuan Wei, Bihui Yu, Wentao Zhang, Bin Cui

机构 * University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) Tsinghua University(清华大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21572 2025-11-14 cs.CL 57%

Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling

Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li

机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09571 2025-11-14 q-bio.QM cs.AI 57%

General Intelligence-based Fragmentation (GIF): A framework for peak-labeled spectra simulation

Margaret R. Martin, Soha Hassoun

机构 * Department of Computer Science, Tufts University, Medford, MA 02155, USA(计算机科学系,塔夫茨大学,马萨诸塞州梅德福,02155,美国)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09539 2025-11-13 cs.CL 57%

SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification

Mohamed Elaraby, Jyoti Prakash Maheswari

机构 * University of Pittsburgh(匹兹堡大学) Zillow Inc.(Zillow公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Journal ref The 5th Workshop on Evaluation & Comparison of NLP Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09313 2025-11-13 cs.CL 57%

Towards Explainable Khmer Polarity Classification

Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09025 2025-11-13 cs.LG 57%

FLAD: Federated Learning for LLM-based Autonomous Driving in Vehicle-Edge-Cloud Networks

Tianao Xiang, Mingjian Zhi, Yuanguo Bi, Lin Cai, Yuhao Chen

机构 * Northeastern University(东北大学) University of Victoria(维多利亚大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08866 2025-11-13 cs.CL 57%

BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation

Fuyi Yang, Chenchen Ye, Mingyu Derek Ma, Yijia Xiao, Matthew Yang, Wei Wang

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24792 2025-11-13 cs.CV cs.AI 57%

PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models

Patrick Haller, Fabio Barth, Jonas Golde, Georg Rehm, Alan Akbik

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 8 pages, 11 tables and figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02615 2025-11-13 astro-ph.IM cs.AI 57%

Radio Astronomy in the Era of Vision-Language Models: Prompt Sensitivity and Adaptation

Mariia Drozdova, Erica Lastufka, Vitaliy Kinakh, Taras Holotyak, Daniel Schaerer, Slava Voloshynovskiy

机构 * University of Geneva(日内瓦大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Machine Learning and the Physical Sciences Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13264 2025-11-13 cs.CL 57%

OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation

Hanning Zhang, Juntong Song, Juno Zhu, Yuanhao Wu, Tong Zhang, Cheng Niu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) NewsBreak

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Preprint update

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08621 2025-11-13 q-fin.ST cs.AI q-fin.CP 57%

The LLM Pro Finance Suite: Multilingual Large Language Models for Financial Applications

Gaëtan Caillaut, Raheel Qader, Jingshu Liu, Mariam Nakhlé, Arezki Sadoune, Massinissa Ahmim, Jean-Gabriel Barthelemy

机构 * Dragon LLM

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01634 2025-11-13 cs.CR cs.AI 57%

Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models

Daniyal Ganiuly, Assel Smaiyl

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08087 2025-11-12 cs.CV cs.AI 57%

Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis

Aditi Singhania, Krutik Malani, Riddhi Dhawan, Arushi Jain, Garv Tandon, Nippun Sharma, Souymodip Chakraborty, Vineet Batra, Ankit Phogat

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08042 2025-11-12 cs.AI 57%

Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations

JV Roig

机构 * Kamiwaza AI

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08028 2025-11-12 cs.LG 57%

Generalizable Insights for Graph Transformers in Theory and Practice

Timo Stoll, Luis Müller, Christopher Morris

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted at NeurIPS 2025 as spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07908 2025-11-12 cs.LG 57%

CellARC: Measuring Intelligence with Cellular Automata

Miroslav Lžičař

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 22 pages, 11 figures. Working draft. Dataset and leaderboard available at https://cellarc.mireklzicar.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07794 2025-11-12 cs.CL 57%

Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark

Hua Zhou, Bing Ma, Yufei Zhang, Yi Zhao

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 16 pages, 11 models,1 set of evaluation framework,5 core dimensions, 54 sub-indicators, 14,430 high-quality questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13861 2025-11-12 cs.HC cs.CL cs.MA 57%

3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark

Ivan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova, Pavel Blinov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) HSE University(俄罗斯高等经济大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments EMNLP 25 (main)

Journal ref https://aclanthology.org/2025.emnlp-main.1353/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07193 2025-11-11 cs.CL 57%

EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language Models

Jiacheng Huang, Ning Yu, Xiaoyin Yi

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06890 2025-11-11 cs.CL 57%

EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers

Yilin Jiang, Mingzi Zhang, Xuanyu Yin, Sheng Jin, Suyu Lu, Zuocan Ying, Zengyi Yu, Xiangjie Kong

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 22 pages, 9 figures, accepted by AAAI2026 as oral paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06738 2025-11-11 cs.CL 57%

Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights

Hyunjae Kim, Jiwoong Sohn, Aidan Gilson, Nicholas Cochran-Caggiano, Serina Applebaum, Heeju Jin, Seihee Park, Yujin Park, Jiyeong Park, Seoyoung Choi, Brittany Alexandra Herrera Contreras, Thomas Huang, Jaehoon Yun, Ethan F. Wei, Roy Jiang, Leah Colucci, Eric Lai, Amisha Dave, Tuo Guo, Maxwell B. Singer, Yonghoe Koo, Ron A. Adelman, James Zou, Andrew Taylor, Arman Cohan, Hua Xu, Qingyu Chen

机构 * Yale School of Medicine(耶鲁医学院) Yale University(耶鲁大学) ETH Zurich(苏黎世联邦理工学院) Harvard Medical School(哈佛医学院) Geisel School of Medicine at Dartmouth(达特茅斯大学盖塞医学院) Seoul National University College of Medicine(首尔国立大学医学院) Hanyang University College of Medicine(翰林大学医学院) PA Leadership Charter School(宾夕法尼亚州西切斯特市领导学校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 34 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04910 2025-11-11 cs.CL 57%

SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents

Jaehoon Lee, Sohyun Kim, Wanggeun Park, Geon Lee, Seungkyung Kim, Minyoung Lee

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 27 pages, 15 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04307 2025-11-11 cs.AI 57%

GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents

Jian Mu, Chaoyun Zhang, Chiming Ni, Lu Wang, Bo Qiao, Kartik Mathur, Qianhui Wu, Yuhang Xie, Xiaojun Ma, Mengyu Zhou, Si Qin, Liqun Li, Yu Kang, Minghua Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

机构 * Nanjing University(南京大学) Microsoft(微软) ZJU-UIUC(浙大-UIUC) Peking University(北京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24826 2025-11-11 cs.CL cs.CV 57%

LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text

Li yunhan, Wu gengshen

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 10 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13949 2025-11-11 cs.CL 57%

Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM

Zheng Wei Lim, Nitish Gupta, Honglin Yu, Trevor Cohn

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05936 2025-11-11 cs.RO cs.AI 57%

10 Open Challenges Steering the Future of Vision-Language-Action Models

Soujanya Poria, Navonil Majumder, Chia-Yu Hung, Amir Ali Bagherzadeh, Chuan Li, Kenneth Kwok, Ziwei Wang, Cheston Tan, Jiajun Wu, David Hsu

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments AAAI 2026 (Senior Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05804 2025-11-11 cs.LG cs.SY eess.SP eess.SY stat.ML 57%

Catching Contamination Before Generation: Spectral Kill Switches for Agents

Valentin Noël

机构 * Devoteam

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Preprint under review (2025). 9 pages, 2 figures. Code and scripts: to be released

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05565 2025-11-11 cs.CV cs.AI 57%

In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy

Shreyan Ganguly, Angona Biswas, Jaydeep Rade, Md Hasibul Hasan Hasib, Nabila Masud, Nitish Singla, Abhipsa Dash, Ushashi Bhattacharjee, Aditya Balu, Anwesha Sarkar, Adarsh Krishnamurthy, Soumik Sarkar

机构 * Iowa State University(爱荷华州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05496 2025-11-11 cs.IR cs.AI 57%

DOCUEVAL: An LLM-based AI Engineering Tool for Building Customisable Document Evaluation Workflows

Hao Zhang, Qinghua Lu, Liming Zhu

机构 * CSIRO’s Data61(CSIRO数据61)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10639 2025-11-11 cs.AR cs.AI 57%

Evaluating LLM-based Workflows for Switched-Mode Power Supply Design

Simon Nau, Jan Krummenauer, André Zimmermann

机构 * Robert Bosch GmbH, Cross-Domain Computing Solutions(罗伯特·博世有限公司,跨领域计算解决方案) University of Stuttgart, Institute for Micro Integration (IFM)(斯图加特大学,微系统集成研究所) Hahn-Schickard, Stuttgart, Germany(哈恩-施克尔德研究所,斯图加特,德国)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏