arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-09-16 至 2025-09-16 共收录 25 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 25 篇

2505.14403 2025-09-16 cs.AI cs.LG 84%

Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning

Zhaohui Yang, Yuxiao Ye, Shilei Jiang, Chen Hu, Linjing Li, Shihong Deng, Daxin Jiang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Institute of Technology(北京理工大学) StepFun, China(中国StepFun)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11648 2025-09-16 cs.CL cs.AI cs.CY 81%

EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI

Sai Kartheek Reddy Kasu

机构 * IIIT Dharwad(德瓦德理工学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10744 2025-09-16 cs.CL cs.AI 81%

Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models

Ozan Gokdemir, Neil Getty, Robert Underwood, Sandeep Madireddy, Franck Cappello, Arvind Ramanathan, Ian T. Foster, Rick L. Stevens

机构 * Argonne National Laboratory(阿贡国家实验室) The University of Chicago(芝加哥大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments This manuscript has been accepted for publication at the Supercomputing 25 (SC '25) Conference (Frontiers in Generative AI for HPC Science and Engineering: Foundations, Challenges, and Opportunities Workshop) in St. Louis, MO, USA on November 16th, 2025. It will appear in the SC25 Workshop Proceedings after that date

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04249 2025-09-16 cs.CL 79%

IOLBENCH: Benchmarking LLMs on Linguistic Reasoning

Satyam Goyal, Soham Dan

机构 * University of Michigan(密歇根大学) Microsoft(微软公司)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11796 2025-09-16 cs.CV 78%

FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning

Haodong Chen, Haojian Huang, XinXiang Yin, Dian Shao

机构 * School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学) The University of Hong Kong(香港大学) School of Software, Northwestern Polytechnical University(软件学院,西北工业大学) Unmanned System Research Institute, Northwestern Polytechnical University(无人系统研究院,西北工业大学)

专题命中 推理评测 :reasoning(title,abstract)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11071 2025-09-16 cs.CV cs.AI cs.CL 73%

The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge

Jinghan Peng, Jingwen Wang, Xing Yu, Dehui Du

机构 * East China Normal University(东华师范大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10345 2025-09-16 cs.CV cs.AI 70%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11026 2025-09-16 cs.AI cs.CL 68%

Rethinking Human Preference Evaluation of LLM Rationales

Ziang Li, Manasi Ganti, Zixian Ma, Helena Vasconcelos, Qijia He, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(abstract,comments);分类 cs.CL、cs.AI;planning(comments)

Comments Published in the XLLM-Reason-Plan Workshop on the Application of LLM Explainability to Reasoning and Planning at COLM 2025

Journal ref Proceedings of the XLLM-Reason-Plan Workshop, Conference on Language Modeling (COLM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11589 2025-09-16 cs.CV 67%

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong, Kaixiang Yang

机构 * Huawei Technologies Co.(华为技术有限公司) South China University of Technology(南方科技大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18240 2025-09-16 cs.CL cs.AI 62%

MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols

Yuhao Du, Qianwei Huang, Guo Zhu, Zhanchen Dai, Shunian Chen, Qiming Zhu, Le Pan, Minghao Chen, Yuhao Zhang, Li Zhou, Benyou Wang, Haizhou Li

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11118 2025-09-16 cs.CL cs.AI 62%

We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism

Priyanshu Priya, Saurav Dudhate, Desai Vishesh Yasheshbhai, Asif Ekbal

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna(计算机科学与工程系,印度理工学院帕纳布)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Paper is accepted at EMNLP (Findings) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00115 2025-09-16 cs.AI cs.CL cs.MA 62%

Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems

Manish Shukla

机构 * Independent Researcher(独立研究者)

专题命中 推理评测 :planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23810 2025-09-16 cs.CL cs.AI 62%

MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation

Chenghao Yang, Yinbo Luo, Zhoufutu Wen, Qi Chu, Tao Gong, Longxiang Liu, Kaiyuan Zhang, Jianpeng Jiao, Ge Zhang, Wenhao Huang, Nenghai Yu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 29 pages, 13 figures, Accepted as EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10584 2025-09-16 cs.CY cs.AI cs.CL 62%

Smart Trial: Evaluating the Use of Large Language Models for Recruiting Clinical Trial Participants via Social Media

Xiaofan Zhou, Zisu Wang, Janice Krieger, Mohan Zalake, Lu Cheng

机构 * University of Illinois at Chicago(伊利诺伊大学香槟分校) Colorado State University(科罗拉多州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12112 2025-09-16 cs.CL 57%

CBP-Tuning: Efficient Local Customization for Black-box Large Language Models

Jiaxuan Zhao, Naibin Gu, Yuchen Feng, Xiyu Liu, Peng Fu, Zheng Lin, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18993 2025-09-16 cs.SE cs.AI 57%

GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging

Ziyi Ni, Huacan Wang, Shuo Zhang, Shuo Lu, Ziyang He, Wang You, Zhenheng Tang, Yuntao Du, Bill Sun, Hongzhang Liu, Sen Hu, Ronghao Chen, Bo Li, Xin Li, Chen Hu, Binxing Jiao, Daxin Jiang, Pin Lyu

机构 * UCAS(中国科学技术大学) CASIA(中国科学院自动化研究所) BUPT(北京理工大学) NUS(新加坡国立大学) StepFun HKUST(香港理工大学) SDU(山东大学) PINAI USYD(悉尼大学) PKU(北京大学) USTC(中国科学技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Highly practical, Well-motivated, Actionable

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16347 2025-09-16 cs.CR cs.AI 57%

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin, Zenghao Duan, zhifei zheng, Min Liu, Zhiyi yin, Jianping Zhang

机构 * State Key Lab of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所状态关键实验室) University of Chinese Academy of Sciences(中国科学院大学) People’s Public Security University of China(中国人民公安大学) Chinese University of Hong Kong(香港大学)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12805 2025-09-16 cs.CL cs.CY cs.HC 57%

Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation

Takaya Arita, Wenxian Zheng, Reiji Suzuki, Fuminori Akiba

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Corrected a typo in the metadata title only ("Assesing"->"Assessing"). No changes were made to the PDF or source files

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18337 2025-09-16 cs.CL 57%

Can LLMs assist with Ambiguity? A Quantitative Evaluation of various Large Language Models on Word Sense Disambiguation

T. G. D. K. Sumanathilaka, Nicholas Micallef, Julian Hough

专题命中 推理评测 :CoT(abstract);分类 cs.CL

Comments 12 pages,6 tables, 1 figure, Proceedings of the 1st International Conference on NLP & AI for Cyber Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10424 2025-09-16 cs.CV cs.AI 57%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10570 2025-09-16 cs.RO cs.AI 57%

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

Wei Dai, Shengen Wu, Wei Wu, Zhenhao Wang, Sisuo Lyu, Haicheng Liao, Limin Yu, Weiping Ding, Runwei Guan, Yutao Yue

机构 * Department of Mathematical Sciences, School of Physical sciences, University of Liverpool(利物浦大学数学科学系) Department of Communications and Networking, School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学通讯与网络系) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) Deep Interdisciplinary Intelligence Lab, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)深度跨学科智能实验室) School of Mathematics and Statistics, Shandong University(山东大学数学与统计学院) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)数据科学与分析方向) Institute of Deep Perception Technology, Jiangsu(江苏深度感知技术研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11952 2025-09-16 cs.CV 50%

CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation

Debopom Sutradhar, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Reem E. Mohamed, Sheikh Izzal Azid, Sami Azam

机构 * Department of Computer Science and Engineering, United International University(计算机科学与工程系,国际大学) Faculty of Science and Information Technology, Charles Darwin University(科学与信息技术学院,查尔斯·达尔文大学) School of Engineering and Energy , Murdoch University(工程与能源学院,默多尼大学) Faculty of Science and Technology, Charles Darwin University(科学与技术学院,查尔斯·达尔文大学)

专题命中 推理评测 :reasoning(abstract)

Comments 23 pages, 6 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11866 2025-09-16 cs.CV 50%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 推理评测 :reasoning(abstract)

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17406 2025-09-16 cs.CV 50%

Seeing the Undefined: Chain-of-Action for Generative Semantic Labels

Meng Wei, Zhongnian Li, Peng Ying, Xinzheng Xu

机构 * 1 School of Computer Science Technology / School of Artificial Intelligence, China University of Mining Technology Xuzhou China 2 Mine Digitization Engineering Research Center of the Ministry of Education Xuzhou China Technology Xuzhou China 2 The State Key Laboratory of CAD\&CG, Zhejiang University Hangzhou China 3 Mine Digitization Engineering Research Center of the Ministry of Education Xuzhou China 2 Mine Digitization Engineering Research Center of the Ministry of Education 2 The State Key Laboratory of CAD\&CG, Zhejiang University 3 Mine Digitization Engineering Research Center of the Ministry of Education

专题命中 推理评测 :reasoning(abstract)

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10563 2025-09-16 cs.CR 50%

Enhancing IoMT Security with Explainable Machine Learning: A Case Study on the CICIOMT2024 Dataset

Mohammed Yacoubi, Omar Moussaoui, C. Drocourt

专题命中 推理评测 :reasoning(abstract)

Journal ref The Third Edition of the International Conference on Connected Objects and Artificial Intelligence (COCIA'2025), Apr 2025, Casablanca (Maroc), Morocco

详情

展开后加载摘要…

URL PDF HTML 收藏