arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10549 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10549 篇

2509.14801 2025-09-19 cs.LG 57%

STEP: Structured Training and Evaluation Platform for benchmarking trajectory prediction models

Julian F. Schumann, Anna Mészáros, Jens Kober, Arkady Zgonnikov

专题命中 推理评测 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14738 2025-09-19 cs.CL 57%

UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

Pengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang, Wei Huang, Zhen Ye, Zhaowei Li, Botian Jiang, Dong Zhang, Xipeng Qiu

机构 * Fudan University(复旦大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by Findings of EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14477 2025-09-19 cs.CL 57%

Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation

Thales Sales Almeida, João Guilherme Alves Santos, Thiago Laitz, Giovana Kerche Bonás

机构 * Institute of Computing (IC) State University of Campinas(计算学院(IC)坎皮纳斯州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18106 2025-09-19 cs.SE cs.AI 57%

A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code

Keke Lian, Bin Wang, Lei Zhang, Libo Chen, Junjie Wang, Ziming Zhao, Yujiu Yang, Miaoqian Lin, Haotong Duan, Haoran Zhao, Shuang Liao, Mingda Guo, Jiazheng Quan, Yilu Zhong, Chenhao He, Zichuan Chen, Jie Wu, Haoling Li, Zhaoxuan Li, Jiongchi Yu, Hui Li, Dong Zhang

机构 * Tencent(腾讯公司) Peking University(北京大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Zhejiang University(浙江大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Singapore Management University(新加坡国立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19477 2025-09-19 cs.AI 57%

Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge

Chiyu Ma, Enpei Zhang, Yilun Zhao, Wenjun Liu, Yaning Jia, Peijun Qing, Lin Shi, Arman Cohan, Yujun Yan, Soroush Vosoughi

机构 * Dartmouth College(达特茅斯学院) Yale University(耶鲁大学)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13773 2025-09-18 cs.AI cs.IR 57%

MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation

Zhipeng Bian, Jieming Zhu, Xuyang Xie, Quanyu Dai, Zhou Zhao, Zhenhua Dong

机构 * Shenzhen University(深圳大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21137 2025-09-18 cs.CL 57%

How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations

Yoshiki Takenami, Yin Jou Huang, Yugo Murawaki, Chenhui Chu

机构 * Kyoto University(京都大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 18 pages, 2 figures. Accepted to EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 57%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05262 2025-09-18 cs.CL 57%

Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmetic

Yang Yan, Yu Lu, Renjun Xu, Zhenzhong Lan

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by EMNLP'25 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13268 2025-09-17 cs.LG 57%

LLMs for energy and macronutrients estimation using only text data from 24-hour dietary recalls: a parameter-efficient fine-tuning experiment using a 10-shot prompt

Rodrigo M Carrillo-Larco

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.LG

Comments https://github.com/rodrigo-carrillo/LLMs-Macronutrient-Estimation-NHANES-Adolescents

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12720 2025-09-17 cs.CL 57%

HistoryBankQA: Multilingual Temporal Question Answering on Historical Events

Biswadip Mandal, Anant Khandelwal, Manish Gupta

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12395 2025-09-17 cs.SE cs.AI 57%

Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML

Yash Mundhra, Max Valk, Maliheh Izadi

机构 * Delft University of Technology(代尔夫特理工大学) ASML(ASML公司)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.AI

Comments Accepted in the 40th IEEE/ACM International Conference on Automated Software Engineering, ASE 2025 (Industry track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09315 2025-09-17 cs.RO cs.CV cs.LG 57%

TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving

Xuefeng Jiang, Yuan Ma, Pengxiang Li, Leimeng Xu, Xin Wen, Kun Zhan, Zhongpu Xia, Peng Jia, Xianpeng Lang, Sheng Sun

机构 * LiAuto Inc(LiAuto公司) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动研究所)

专题命中 推理评测 :planning(abstract);分类 cs.LG

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22388 2025-09-17 cs.CL 57%

Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors

Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng

机构 * Singapore Management University(新加坡管理大学) Tsinghua University(清华大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Main, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12112 2025-09-16 cs.CL 57%

CBP-Tuning: Efficient Local Customization for Black-box Large Language Models

Jiaxuan Zhao, Naibin Gu, Yuchen Feng, Xiyu Liu, Peng Fu, Zheng Lin, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18993 2025-09-16 cs.SE cs.AI 57%

GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging

Ziyi Ni, Huacan Wang, Shuo Zhang, Shuo Lu, Ziyang He, Wang You, Zhenheng Tang, Yuntao Du, Bill Sun, Hongzhang Liu, Sen Hu, Ronghao Chen, Bo Li, Xin Li, Chen Hu, Binxing Jiao, Daxin Jiang, Pin Lyu

机构 * UCAS(中国科学技术大学) CASIA(中国科学院自动化研究所) BUPT(北京理工大学) NUS(新加坡国立大学) StepFun HKUST(香港理工大学) SDU(山东大学) PINAI USYD(悉尼大学) PKU(北京大学) USTC(中国科学技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Highly practical, Well-motivated, Actionable

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16347 2025-09-16 cs.CR cs.AI 57%

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin, Zenghao Duan, zhifei zheng, Min Liu, Zhiyi yin, Jianping Zhang

机构 * State Key Lab of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所状态关键实验室) University of Chinese Academy of Sciences(中国科学院大学) People’s Public Security University of China(中国人民公安大学) Chinese University of Hong Kong(香港大学)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12805 2025-09-16 cs.CL cs.CY cs.HC 57%

Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation

Takaya Arita, Wenxian Zheng, Reiji Suzuki, Fuminori Akiba

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Corrected a typo in the metadata title only ("Assesing"->"Assessing"). No changes were made to the PDF or source files

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18337 2025-09-16 cs.CL 57%

Can LLMs assist with Ambiguity? A Quantitative Evaluation of various Large Language Models on Word Sense Disambiguation

T. G. D. K. Sumanathilaka, Nicholas Micallef, Julian Hough

专题命中 推理评测 :CoT(abstract);分类 cs.CL

Comments 12 pages,6 tables, 1 figure, Proceedings of the 1st International Conference on NLP & AI for Cyber Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10424 2025-09-16 cs.CV cs.AI 57%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10570 2025-09-16 cs.RO cs.AI 57%

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

Wei Dai, Shengen Wu, Wei Wu, Zhenhao Wang, Sisuo Lyu, Haicheng Liao, Limin Yu, Weiping Ding, Runwei Guan, Yutao Yue

机构 * Department of Mathematical Sciences, School of Physical sciences, University of Liverpool(利物浦大学数学科学系) Department of Communications and Networking, School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学通讯与网络系) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) Deep Interdisciplinary Intelligence Lab, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)深度跨学科智能实验室) School of Mathematics and Statistics, Shandong University(山东大学数学与统计学院) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)数据科学与分析方向) Institute of Deep Perception Technology, Jiangsu(江苏深度感知技术研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13335 2025-09-15 cs.CL 57%

Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes

Tyler Loakman, William Thorne, Chenghua Lin

机构 * Department of Computer Science, The University of Sheffield, UK(谢菲尔德大学计算机科学系) Department of Computer Science, The University of Manchester, UK(曼彻斯特大学计算机科学系)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00858 2025-09-15 cs.AI cs.HC 57%

Learning to Plan with Personalized Preferences

Manjie Xu, Xinyi Yang, Wei Liang, Chi Zhang, Yixin Zhu

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) School of Computer Science & Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) Yangtze Delta Region Academy of Beijing Institute of Technology, Jiaxing, China(北京理工大学扬子江地区研究院) National Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09614 2025-09-12 cs.SE cs.AI 57%

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Jielin Qiu, Zuxin Liu, Zhiwei Liu, Rithesh Murthy, Jianguo Zhang, Haolin Chen, Shiyu Wang, Ming Zhu, Liangwei Yang, Juntao Tan, Zhepeng Cen, Cheng Qian, Shelby Heinecke, Weiran Yao, Silvio Savarese, Caiming Xiong, Huan Wang

机构 * Salesforce AI Research(Salesforce AI研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 53 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01523 2025-09-12 cs.CL 57%

CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering

Zongxi Li, Yang Li, Haoran Xie, S. Joe Qin

机构 * School of Data Science, Lingnan University(数据科学学院) School of Science and Technology, Hong Kong Metropolitan University(科技学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08008 2025-09-11 cs.SI cs.AI cs.MM 57%

A New Dataset and Benchmark for Grounding Multimodal Misinformation

Bingjian Yang, Danni Xu, Kaipeng Niu, Wenxuan Liu, Zheng Wang, Mohan Kankanhalli

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(多媒体软件国家工程研究中心,计算机科学学院,武汉大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学) School of Computer Science, Peking University(计算机科学学院,北京大学) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 5 figures, ACM Multimedia 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07817 2025-09-10 cs.CL cs.MM 57%

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * National University of Singapore(新加坡国立大学) Southern University of Science and Technology(南方科技大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07190 2025-09-10 cs.CL cs.HC 57%

Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation

Zahra Atf, Peter R Lewis

机构 * Faculty of Business and Information Technology(商业与信息技术学院) Ontario Tech University(安大略技术大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments This paper was accepted for presentation at the 35th IEEE International Conference on Collaborative Advances in Software and Computing. Conference website:https://conf.researchr.org/home/cascon-2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07135 2025-09-10 cs.CL 57%

MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations

Ruggero Marino Lazzaroni, Alessandro Angioi, Michelangelo Puliga, Davide Sanna, Roberto Marras

机构 * University of Graz(格拉茨大学) OnePix Academy(OnePix学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted as an oral presentation at CLiC-it 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04656 2025-09-10 cs.CL 57%

AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

Aisha Alansari, Hamzah Luqman

机构 * Information and Computer Science Department, King Fahd University of Petroleum and Minerals(信息与计算机科学系,国王法赫德石油和矿物大学) SDAIA-KFUPM Joint Research Center for Artificial Intelligence(SDAIA-KFUPM人工智能联合研究中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏