arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10577 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10577 篇

2506.06133 2025-06-09 cs.CL 57%

Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition

Tara Azin, Daniel Dumitrescu, Diana Inkpen, Raj Singh

机构 * Carleton University(卡尔顿大学) University of Ottawa(渥太华大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments This paper is published in the Proceedings of the 38th Canadian Conference on Artificial Intelligence (CAIAC 2025). Please cite the conference version at https://caiac.pubpub.org/pub/keh8ij01

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06093 2025-06-09 cs.CL 57%

Reinforcing Code Generation: Improving Text-to-SQL with Execution-Based Learning

Atharv Kulkarni, Vivek Srikumar

机构 * University of Utah(犹他大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Under review at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05817 2025-06-09 cs.SE cs.CL 57%

CodeContests+: High-Quality Test Case Generation for Competitive Programming

Zihan Wang, Siyao Liu, Yang Sun, Hongyan Li, Kai Shen

机构 * ByteDance Seed(字节跳动种子) Peking University(北京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 28 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00814 2025-06-09 cs.CL 57%

GuessBench: Sensemaking Multimodal Creativity in the Wild

Zifeng Zhu, Shangbin Feng, Herun Wan, Ningnan Wang, Minnan Luo, Yulia Tsvetkov

机构 * Xi’an Jiaotong University(西安交通大学) University of Washington(华盛顿大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04756 2025-06-06 cs.AI cs.CV eess.IV 57%

Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems

Loan Dao, Ngoc Quoc Ly

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04454 2025-06-06 cs.LG 57%

Neurosymbolic Artificial Intelligence for Robust Network Intrusion Detection: From Scratch to Transfer Learning

Huynh T. T. Tran, Jacob Sander, Achraf Cohen, Brian Jalaian, Nathaniel D. Bastian

机构 * University of West Florida(西弗吉尼亚大学) United States Military Academy(美国军事学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 17 pages, 5 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17213 2025-06-06 cs.CV cs.AI 57%

VCD: A Dataset for Visual Commonsense Discovery in Images

Xiangqing Shen, Fanfan Wang, Siwei Wu, Rui Xia

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03510 2025-06-05 cs.CL 57%

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information

Seungcheol Park, Sojin Lee, Jongjin Kim, Jinsik Lee, Hyunjik Jo, U Kang

机构 * Seoul National University(首尔国立大学) LG AI Research(LG AI研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments IJCAI 2025 Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03332 2025-06-05 cs.AI 57%

Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows

Yifei Ming, Zixuan Ke, Xuan-Phi Nguyen, Jiayu Wang, Shafiq Joty

机构 * Salesforce AI Research(Salesforce人工智能研究) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03278 2025-06-05 cs.CL 57%

FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes

Christodoulos Constantinides, Dhaval Patel, Shuxin Lin, Claudio Guerrero, Sunil Dagajirao Patil, Jayant Kalagnanam

机构 * IBM TJ Watson Research Center(IBM TJ Watson 研究中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00264 2025-06-05 cs.CL 57%

MultiHoax: A Dataset of Multi-hop False-Premise Questions

Mohammadamin Shafiei, Hamidreza Saffari, Nafise Sadat Moosavi

机构 * University of Milan(米兰大学) Politecnico di Milano(米兰理工学院) University of Sheffield(谢菲尔德大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments accepted at ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23899 2025-06-05 cs.CL 57%

Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset

Diana Galvan-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi gu, Zheng Yuan, Keisuke Sakaguchi, Paula Buttery

机构 * ALTA Institute, Computer Laboratory, University of Cambridge(ALTA研究所、计算机实验室、剑桥大学) SB Intuitions Tohoku University(东北大学) RIKEN(日本理化学研究所) The University of Sheffield(谢菲尔德大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 10 main pages (24 appendix pages), 9 figures, accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09887 2025-06-05 cs.LG math.OC 57%

OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling

Zhicheng Yang, Yiwei Wang, Yinya Huang, Zhijiang Guo, Wei Shi, Xiongwei Han, Liang Feng, Linqi Song, Xiaodan Liang, Jing Tang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of California, Merced(加州大学梅德福分校) ETH Zurich(苏黎世联邦理工学院) City University of Hong Kong(香港城市大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Sun Yat-sen University(中山大学) MBZUAI Chongqing University(重庆大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Journal ref The Thirteenth International Conference on Learning Representations, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03139 2025-06-04 cs.CV cs.AI 57%

SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation

Siqi Chen, Xinyu Dong, Haolei Xu, Xingyu Wu, Fei Tang, Hang Zhang, Yuchen Yan, Linjuan Wu, Wenqi Zhang, Guiyang Hou, Yongliang Shen, Weiming Lu, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 19 pages,4 figures, Project page: https://zju-real.github.io/SVGenius, Code: https://github.com/ZJU-REAL/SVGenius-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03090 2025-06-04 cs.CL 57%

Literary Evidence Retrieval via Long-Context Language Models

Katherine Thai, Mohit Iyyer

机构 * UMass Amherst(马萨诸塞大学阿姆赫斯特分校) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02167 2025-06-04 cs.CV cs.AI 57%

Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360-Degree Firefighting Videos

Aditi Tiwari, Farzaneh Masoud, Dac Trong Nguyen, Jill Kraft, Heng Ji, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Illinois Fire Service Institute(伊利诺伊州消防服务研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 20 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02046 2025-06-04 cs.CY cs.AI 57%

Machine vs Machine: Using AI to Tackle Generative AI Threats in Assessment

Mohammad Saleh Torkestani, Taha Mansouri

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Paper presented at the Learning, Teaching & Student Experience 2025 Conference. The Chartered Association of Business Schools (CABS), Nottingham, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03730 2025-06-04 cs.LG cs.CR cs.CV 57%

NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA

Marlon Tobaben, Mohamed Ali Souibgui, Rubèn Tito, Khanh Nguyen, Raouf Kerkouche, Kangsoo Jung, Joonas Jälkö, Lei Kang, Andrey Barsky, Vincent Poulain d'Andecy, Aurélie Joseph, Aashiq Muhamed, Kevin Kuo, Virginia Smith, Yusuke Yamasaki, Takumi Fukami, Kenta Niwa, Iifan Tyou, Hiro Ishii, Rio Yokota, Ragul N, Rintu Kutum, Josep Llados, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas

机构 * University of Helsinki(赫尔辛基大学) Computer Vision Center, Universitat Autònoma de Barcelona(巴塞罗那自治大学计算机视觉中心) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍茨中心) INRIA(法国国家信息与自动化技术研究所) Yooz Carnegie Mellon University(卡内基梅隆大学) NTT(日本NTT公司) Institute of Science Tokyo(东京科学研究所) Department of Computer Science(计算机科学系) Mphasis AI & Applied Tech Lab at Ashoka, Ashoka University(阿什oka大学人工智能与应用技术实验室) Koita Centre for Digital Health - Ashoka (KCDH-A)(阿什oka数字健康中心(KCDH-A)) Trivedi School of Biosciences, Ashoka University(阿什oka大学Trivedi生物科学学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 33 pages, 7 figures; published in TMLR 06/2025 https://openreview.net/forum?id=3HKNwejEEq

Journal ref Transactions on Machine Learning Research, ISSN 2835-8856, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03525 2025-06-04 cs.CL 57%

UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization

Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven R. Corman, Chitta Baral

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted at ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01937 2025-06-03 cs.CL 57%

RewardBench 2: Advancing Reward Model Evaluation

Saumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison, Noah A. Smith, Hannaneh Hajishirzi, Nathan Lambert

机构 * Allen Institute for Artificial Intelligence(艾伦人工智能研究所) University of Washington(华盛顿大学) Cohere

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Data, models, and leaderboard available at https://huggingface.co/collections/allenai/reward-bench-2-683d2612a4b3e38a3e53bb51

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01814 2025-06-03 cs.CL cs.SI 57%

Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high

PeiHsuan Huang, ZihWei Lin, Simon Imbot, WenCheng Fu, Ethan Tu

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01676 2025-06-03 cs.AI 57%

K12Vista: Exploring the Boundaries of MLLMs in K-12 Education

Chong Li, Chenglin Zhu, Tao Zhang, Mingan Lin, Zenan Zhou, Jian Xie

机构 * Baichuan Inc(百川公司) Peking University(北京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01520 2025-06-03 cs.CL 57%

FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents

Bobo Li, Yuheng Wang, Hao Fei, Juncheng Li, Wei Ji, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) Wuhan University(武汉大学) Zhejiang University(浙江大学) Nanjing University(南京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01344 2025-06-03 cs.CL 57%

Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents

Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Vivek Gupta, Dinesh Manocha

机构 * University of Maryland(马里兰大学) Adobe Research(Adobe研究) ASU(亚利桑那州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00855 2025-06-03 cs.AI 57%

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

Sau Lai Yip, Sunan He, Yuxiang Nie, Shu Pui Chan, Yilin Ye, Sum Ying Lam, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(化学与生物工程系,香港科学与技术大学) Division of Life Science, The Hong Kong University of Science and Technology(生命科学系,香港科学与技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments For data and code, see: https://huggingface.co/datasets/slyipae1/MedBookVQA and https://github.com/slyipae1/MedBookVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19433 2025-06-03 cs.LG 57%

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression

Peijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li, Xiaowen Chu, Bo Li

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted by ICML2025 as Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00212 2025-06-03 cs.MA cs.CL 57%

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang, Huazheng Wang, Yiran Chen, Qingyun Wu

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Duke University(杜克大学) University of Washington(华盛顿大学) Oregon State University(俄勒冈州立大学) Google DeepMind(谷歌DeepMind) Nanyang Technological University(南洋理工大学) Meta AG2AI, Inc.(AG2AI公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16167 2025-06-03 cs.SE cs.CL 57%

CodeReviewQA: The Code Review Comprehension Assessment for Large Language Models

Hong Yi Lin, Chunhua Liu, Haoyu Gao, Patanamon Thongtanunam, Christoph Treude

机构 * The University of Melbourne(墨尔本大学) Singapore Management University(新加坡管理学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments The paper is published in Findings of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20584 2025-06-03 cs.CL 57%

Towards Neural No-Resource Language Translation: A Comparative Evaluation of Approaches

Madhavendra Thakur

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Presented at the Columbia AI Summit 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17451 2025-06-03 cs.CV cs.CL 57%

VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Lei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang, Yifan Song, Peiyi Wang, Chenxin An, Tianyu Liu, Sujian Li, Bill Yuchen Lin, Lingpeng Kong, Qi Liu

机构 * HKU(香港大学) SCUT(华南理工大学) SJTU(上海交通大学) PKU(北京大学) Allen AI(AllenAI)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments CVPR 2025 Camera Ready Version. Project page: https://vl-rewardbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏