arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10549 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10549 篇

2510.07363 2025-10-15 cs.AI 74%

L2M-AID: Autonomous Cyber-Physical Defense by Fusing Semantic Reasoning of Large Language Models with Multi-Agent Reinforcement Learning (Preprint)

Tianxiang Xu, Zhichao Wen, Xinyu Zhao, Jun Wang, Yan Li, Chang Liu

机构 * Peking University(北京大学) RWTH Aachen University(亚琛工业大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Wuhan University(武汉大学) Thales Group(泰雷兹集团) Chinese Medical Information and Big Data Association(中国医学信息与大数据协会)

专题命中 推理评测 :reasoning(title);分类 cs.AI

Comments This preprint was submitted to IEEE TrustCom 2025. The accepted version will be published under copyright 2025 IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00962 2025-10-02 cs.CL 74%

Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks

Eileen Pan, Anna Seo Gyeong Choi, Maartje ter Hoeve, Skyler Seto, Allison Koenecke

机构 * Department of Information Science, Cornell University(康奈尔大学信息科学系) Apple(苹果公司) Cornell Tech(康奈尔科技)

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments EMNLP Findings 2025, 12 pages, 11 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02604 2025-09-26 cs.LG stat.ME 74%

Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables

Ivaxi Sheth, Sahar Abdelnabi, Mario Fritz

机构 * CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全中心) Microsoft(微软公司)

专题命中 推理评测 :reasoning(title);分类 cs.LG

Comments EMNLP'25 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13723 2025-09-19 cs.CL 74%

DSPC: Dual-Stage Progressive Compression Framework for Efficient Long-Context Reasoning

Yaxin Gao, Yao Lu, Zongfei Zhang, Jiaqi Nie, Shanqing Yu, Qi Xuan

机构 * Institute of Cyberspace Security, Zhejiang University of Technology A STAR Amazon Binjiang Institute of Artificial Intelligence, Zhejiang University of Technology

专题命中 推理评测 :reasoning(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21763 2025-07-22 cs.AI 74%

THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?

Xin Wang, Jiyao Liu, Yulong Xiao, Junzhi Ning, Lihao Liu, Junjun He, Botian Shi, Kaicheng Yu

机构 * Westlake University(西湖大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Fuzhou University(福州市大学) Imperial College London(伦敦帝国学院)

专题命中 推理评测 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24714 2025-06-02 cs.CL 74%

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation

Junyu Luo, Zhizhuo Kou, Liming Yang, Xiao Luo, Jinsheng Huang, Zhiping Xiao, Jingshu Peng, Chengzhong Liu, Jiaming Ji, Xuanzhe Liu, Sirui Han, Ming Zhang, Yike Guo

机构 * State Key Laboratory for Multimedia Information Processing, PKU-Anker LLM Lab(多媒体信息处理国家重点实验室,PKU-Anker LLM实验室) School of Computer Science, Peking University(北京大学计算机学院) HKUST(香港科技大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Washington(华盛顿大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18291 2025-05-28 cs.CV cs.CL cs.RO 74%

InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning

Zifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin, Zihan Wang, Simon Stepputtis, Deva Ramanan, Katia Sycara

机构 * Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments Accepted by ACL 2025 Main. Project page: https://zifuwan.github.io/InstructPart/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18341 2025-05-27 cs.RO cs.AI 74%

CrashAgent: Crash Scenario Generation via Multi-modal Reasoning

Miao Li, Wenhao Ding, Haohong Lin, Yiqi Lyu, Yihang Yao, Yuyou Zhang, Ding Zhao

机构 * Carnegie Mellon University(卡内基梅隆大学) NVIDIA Research(NVIDIA研究) Northwestern University(西北大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08498 2025-05-26 cs.CL 74%

"Reasoning" with Rhetoric: On the Style-Evidence Tradeoff in LLM-Generated Counter-Arguments

Preetika Verma, Kokil Jaidka, Svetlana Churina

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments 24 pages, 9 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03970 2025-05-08 cs.CL 74%

A Reasoning-Focused Legal Retrieval Benchmark

Lucia Zheng, Neel Guha, Javokhir Arifov, Sarah Zhang, Michal Skreta, Christopher D. Manning, Peter Henderson, Daniel E. Ho

机构 * Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments CS&Law 2025. For data, see https://reglab.github.io/legal-rag-benchmarks/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13143 2025-04-18 cs.CV cs.AI 74%

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

Siwei Yang, Mude Hui, Bingchen Zhao, Yuyin Zhou, Nataniel Ruiz, Cihang Xie

专题命中 推理评测 :CoT(title);分类 cs.AI

Comments Project Page: https://ucsc-vlaa.github.io/Complex-Edit/, Dataset: https://huggingface.co/datasets/UCSC-VLAA/Complex-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22152 2025-03-31 cs.CV cs.AI 74%

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

Yuxuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads, Asli Celikyilmaz, Karl Ridgeway

专题命中 推理评测 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04726 2025-03-28 cs.CL cs.CV 74%

Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation

Jeff Da, Maxwell Forbes, Rowan Zellers, Anthony Zheng, Jena D. Hwang, Antoine Bosselut, Yejin Choi

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments ACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15478 2025-03-20 cs.LG 74%

SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Yifei Zhou, Song Jiang, Yuandong Tian, Jason Weston, Sergey Levine, Sainbayar Sukhbaatar, Xian Li

专题命中 推理评测 :reasoning(title);分类 cs.LG

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06396 2024-10-10 cs.CL 74%

MLissard: Multilingual Long and Simple Sequential Reasoning Benchmarks

Mirelle Bueno, Roberto Lotufo, Rodrigo Nogueira

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments GenBench Workshop by EMNLP 2024: Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04818 2024-09-04 cs.CL 74%

ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi, Kentaro Inui

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments 18 pages, 7 figures, accepted to COLM 2024. Data available here: https://github.com/a-brassard/ACORN

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00126 2024-07-19 cs.CV cs.CL 74%

Common Sense Reasoning for Deepfake Detection

Yue Zhang, Ben Colman, Xiao Guo, Ali Shahriyari, Gaurav Bharaj

专题命中 推理评测 :reasoning(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16490 2024-06-25 cs.CL 74%

eagerlearners at SemEval2024 Task 5: The Legal Argument Reasoning Task in Civil Procedure

Hoorieh Sabzevari, Mohammadmostafa Rostamkhani, Sauleh Eetemadi

专题命中 推理评测 :reasoning(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00822 2024-06-12 cs.AI cs.HC cs.RO 74%

Open-Ended Multi-Modal Relational Reasoning for Video Question Answering

Haozheng Luo, Ruiyang Qin, Chenwei Xu, Guo Ye, Zening Luo

专题命中 推理评测 :reasoning(title);分类 cs.AI

Comments 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01573 2024-06-06 cs.SE cs.AI 74%

Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository

Ajinkya Deshpande, Anmol Agarwal, Shashank Shet, Arun Iyer, Aditya Kanade, Ramakrishna Bairi, Suresh Parthasarathy

专题命中 推理评测 :reasoning(title);分类 cs.AI

Comments Preprint with additional experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12541 2024-05-22 cs.AI 74%

DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge

Bufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu, Hai Li, Guoliang Xing, Hongkai Chen, Xiaofan Jiang, Zhenyu Yan

专题命中 推理评测 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15737 2024-03-26 cs.CL 74%

Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning

Zhouhang Xie, Bodhisattwa Prasad Majumder, Mengjie Zhao, Yoshinori Maeda, Keiichi Yamada, Hiromi Wakaki, Julian McAuley

专题命中 推理评测 :reasoning(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04706 2024-01-01 cs.LG 74%

Offline Imitation Learning with Variational Counterfactual Reasoning

Bowei He, Zexu Sun, Jinxin Liu, Shuai Zhang, Xu Chen, Chen Ma

专题命中 推理评测 :reasoning(title);分类 cs.LG

Comments Published on NeurIPS2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01684 2023-10-24 cs.CL 74%

Baby's CoThought: Leveraging Large Language Models for Enhanced Reasoning in Compact Models

Zheyu Zhang, Han Yang, Bolei Ma, David Rügamer, Ercong Nie

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments CoNLL 2023 BabyLM Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02258 2023-06-06 cs.CL 74%

Probing Physical Reasoning with Counter-Commonsense Context

Kazushi Kondo, Saku Sugawara, Akiko Aizawa

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments Accepted to ACL 2023(Short Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14763 2023-05-25 cs.CL 74%

Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, Vered Shwartz

专题命中 推理评测 :reasoning(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07269 2023-02-17 cs.CL 74%

SODAPOP: Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models

Haozhe An, Zongxia Li, Jieyu Zhao, Rachel Rudinger

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02950 2022-11-08 cs.CL 74%

The Legal Argument Reasoning Task in Civil Procedure

Leonard Bongard, Lena Held, Ivan Habernal

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments Camera ready, to appear at the Natural Legal Language Processing Workshop 2022 co-located with EMNLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16952 2022-11-07 cs.CL 74%

Transfer Learning with Synthetic Corpora for Spatial Role Labeling and Reasoning

Roshanak Mirzaee, Parisa Kordjamshidi

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15109 2022-10-03 cs.CL 74%

ConceptNet infused DialoGPT for Underlying Commonsense Understanding and Reasoning in Dialogue Response Generation

Ye Liu, Wolfgang Maier, Wolfgang Minker, Stefan Ultes

专题命中 推理评测 :reasoning(title);分类 cs.CL

Comments this is a long paper, the short version was accepted by SemDial 2022

详情

展开后加载摘要…

URL PDF HTML 收藏