arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-30 至 2025-10-30 共收录 14 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 14 篇

2510.25332 2025-10-30 cs.CV 89%

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu

机构 * Henan Institute of Advanced Technology, Zhengzhou University(河南高级技术研究所,郑州大学) Institute of Automation, CAS(自动化研究所,中国科学院) UCAS(中国科学院大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24932 2025-10-30 cs.CL 83%

RiddleBench: A New Generative Reasoning Benchmark for LLMs

Deepon Halder, Alan Saji, Thanmay Jayakumar, Ratish Puduppully, Anoop Kunchukuttan, Raj Dabre

机构 * Nilekani Centre at AI4Bharat(AI4Bharat的Nilekani中心) Indian Institute of Technology Madras(印度理工学院马德拉斯分校) IT University of Copenhagen(哥本哈根技术大学) Microsoft(微软) Indian Institute of Engineering, Science and Technology, Shibpur(Shibpur印度工程科学技术学院)

专题命中 推理评测 :reasoning(title,abstract);self-correction(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04458 2025-10-30 cs.CL 83%

Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection

Soumyadeep Jana, Abhrajyoti Kundu, Sanasam Ranbir Singh

机构 * Indian Institute of Technology Guwahati(印度理工学院古瓦哈提)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25187 2025-10-30 cs.CL 70%

Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction

Ritesh Sunil Chavan, Jack Mostow

机构 * Department of Computer Science(计算机科学系) Stony Brook University(石溪大学) School of Computer Science(计算机科学学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15584 2025-10-30 cs.HC 67%

To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models

Jessica Y. Bo, Sophia Wan, Ashton Anderson

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract)

Journal ref Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09135 2025-10-30 cs.AI cs.CL cs.HC cs.LG 67%

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L. Holt, Fernando De la Torre

机构 * Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所) Department of Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校心理学系) Center for Perceptual Systems, The University of Texas at Austin(德克萨斯大学奥斯汀分校感知系统中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 22 pages, first three authors equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06677 2025-10-30 cs.RO cs.CV 67%

RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation

Songhao Han, Boxiang Qiu, Yue Liao, Siyuan Huang, Chen Gao, Shuicheng Yan, Si Liu

机构 * Beihang University(北航大学) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract)

Comments 25 pages, 18 figures, Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25694 2025-10-30 cs.SE cs.AI cs.CL 62%

Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents

Jiayi Kuang, Yinghui Li, Xin Zhang, Yangning Li, Di Yin, Xing Sun, Ying Shen, Philip S. Yu

机构 * Youtu-LLM Team, Tencent Youtu Lab(腾讯优设实验室) Sun Yat-sen University(中山大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 推理评测 :planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24788 2025-10-30 cs.CV cs.AI cs.LG 62%

The Underappreciated Power of Vision Models for Graph Structural Understanding

Xinjian Zhao, Wei Pang, Zhongkai Xue, Xiangru Jian, Lei Zhang, Yaoyao Xu, Xiaozhuang Song, Shu Wu, Tianshu Yu

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Cheriton School of Computer Science, University of Waterloo(滑铁卢大学切尔顿计算机科学学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24762 2025-10-30 cs.CL cs.AI 62%

Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation

Wenzhen Luo, Wei Guan, Yifan Yao, Yimin Pan, Feng Wang, Zhipeng Yu, Zhe Wen, Liang Chen, Yihong Zhuang

机构 * Ant Group(蚂蚁集团)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22967 2025-10-30 cs.CL cs.AI 62%

MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs

Yucheng Ning, Xixun Lin, Fang Fang, Yanan Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100085, China(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100049, China(中国科学院大学网络安全学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-025-51369-x}

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10844 2025-10-30 cs.AI cs.CL 62%

Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models

Simeng Han, Howard Dai, Stephen Xia, Grant Zhang, Chen Liu, Lichang Chen, Hoang Huy Nguyen, Hongyuan Mei, Jiayuan Mao, R. Thomas McCoy

机构 * Yale University(耶鲁大学) Meta Superintelligence Labs(Meta超智能实验室) Georgia Institute of Technology(佐治亚理工学院) TTIC Massachusetts Institute of Technology(麻省理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.06818 2025-10-30 cs.CL cs.IR 57%

Large Language Models for Few-Shot Named Entity Recognition

Yufei Zhao, Xiaoshi Zhong, Erik Cambria, Jagath C. Rajapakse

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL

Comments 17 pages, 2 figures. Accepted by AI, Computer Science and Robotics Technology (ACRT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19028 2025-10-30 cs.CV cs.AI 57%

InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

Tianchi Xie, Minzhi Lin, Mengchen Liu, Yilin Ye, Changjian Chen, Shixia Liu

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏