arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10530 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10530 篇

2504.20781 2025-12-10 cs.SE cs.AI 70%

Using LLMs in Generating Design Rationale for Software Architecture Decisions

在软件架构决策中使用LLMs生成设计理由

Xiyu Zhou, Ruiyin Li, Peng Liang, Beiqi Zhang, Mojtaba Shahin, Zengyang Li, Chen Yang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) RMIT University(皇家墨尔本理工大学) School of Computer Science, Central China Normal University(中央师范大学计算机学院) School of Artificial Intelligence, Shenzhen Polytechnic University(深圳职业技术学院人工智能学院)

专题命中 推理评测 :reasoning(abstract);CoT(abstract);分类 cs.AI

AI总结 本研究评估了LLMs在生成软件架构决策设计理由方面的性能,通过实验和访谈探讨了不同提示策略的效果及实际应用的可行性。

Comments Preprint accepted for publication in ACM Transactions on Software Engineering and Methodology (TOSEM), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05339 2025-12-08 cs.LG 70%

Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models

具有鲁棒防护栏的分类适应调谐模型用于大型语言模型

Mahesh Kumar Nandwana, Youngwan Lim, Joseph Liu, Alex Yang, Varun Notibala, Nishchaie Khanna

机构 * Project Lead(项目负责人)

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.LG

AI总结 Roblox Guard 1.0通过指令微调提升大型语言模型的安全性,采用输入输出调谐和可扩展的安全分类评估框架。

Comments To be presented at AAAI-26 PerFM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03558 2025-12-04 cs.CV cs.CL 70%

CartoMapQA: A Fundamental Benchmark Dataset Evaluating Vision-Language Models on Cartographic Map Understanding

CartoMapQA: 一个评估视觉-语言模型在制图地图理解上的基础基准数据集

Huy Quang Ung, Guillaume Habault, Yasutaka Nishimura, Hao Niu, Roberto Legaspi, Tomoki Oya, Ryoichi Kojima, Masato Taya, Chihiro Ono, Atsunori Minamikawa, Yan Liu

机构 * KDDI Research, Inc.(KDDI研究公司) University of Southern California(南加州大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.CL

AI总结 CartoMapQA通过问答任务评估视觉-语言模型在制图地图理解上的能力,揭示了模型在地图语义和地理推理方面的不足。

Comments Accepted at SIGSPATIAL 2025 (Best paper candidates), 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21444 2025-11-27 cs.AI physics.ao-ph 70%

EWE: An Agentic Framework for Extreme Weather Analysis

EWE:极端天气分析的代理框架

Zhe Jiang, Jiong Wang, Xiaoyu Yue, Zijie Guo, Wenlong Zhang, Fenghua Ling, Wanli Ouyang, Lei Bai

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The University of Sydney(悉尼大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

AI总结 EWE是首个用于极端天气分析的智能代理框架,通过知识引导的规划和闭环推理实现自动化诊断,提供首个该领域基准测试,推动科学发现民主化进程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16383 2025-11-26 cs.CL 70%

Large Language Models in Argument Mining: A Survey

大型语言模型在论证挖掘中的应用:综述

Hao Li, Viktor Schlegel, Yizheng Sun, Riza Batista-Navarro, Goran Nenadic

机构 * University of Manchester(曼彻斯特大学) Imperial College London(伦敦帝国学院)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 本文综述了大型语言模型对论证挖掘领域的影响,分析了LLM如何改变任务设计、数据集构建和评估方法,并提出了未来研究方向。

Comments Work draft

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18298 2025-11-25 cs.AI 70%

Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery

跨学科知识检索与综合:一种用于科学发现的复合AI架构

Svitlana Volkova, Peter Bautista, Avinash Hiriyanna, Gabriel Ganberg, Isabel Erickson, Zachary Klinefelter, Nick Abele, Hsien-Te Kao, Grant Engberson

机构 * Aptima, Inc.(Aptima公司)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

AI总结 BioSage通过整合LLMs与RAG,利用专门代理实现跨学科知识检索与综合,提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15511 2025-11-24 cs.RO cs.AI 70%

AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models

AeroVerse:用于模拟、预训练、微调和评估航空航天具身世界模型的UAV-Agent基准套件

Fanglong Yao, Yuanchang Yue, Youzhi Liu, Xian Sun, Kun Fu

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) Key Laboratory of Target Cognition and Application Technology(TCAT), Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所目标认知与应用技术重点实验室)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

AI总结 AeroVerse是一个用于模拟、预训练、微调和评估航空航天具身世界模型的基准套件,包含多个数据集和评估指标,旨在推动航空航天具身智能的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16221 2025-11-21 cs.CV cs.CL 70%

Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions

大语言模型能读取环境吗?一个多模态基准用于评估多当事人社交互动中的欺骗

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Ruicong Liu, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 本文提出多模态互动欺骗评估任务,通过新颖数据集评估多种MLLMs的欺骗检测能力,揭示其在多模态社交线索处理上的不足,并提出SoCoT和DSEM模块提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11252 2025-11-17 cs.AI 70%

UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios

Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

机构 * Department of Computer and Network Engineering, College of Information Technology, United Arab Emirates University(计算机与网络工程系,信息科技学院,阿联酋大学) Khalifa University of Science and Technology(科技大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

Comments 18 pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10289 2025-11-14 eess.AS cs.CL 70%

Music Flamingo: Scaling Music Understanding in Audio Language Models

Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee, Zhifeng Kong, Joao Felipe Santos, Ramani Duraiswami, Dinesh Manocha, Wei Ping, Mohammad Shoeybi, Bryan Catanzaro

机构 * NVIDIA, CA, USA(NVIDIA公司) University of Maryland, College Park, USA(大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments Project Page: https://research.nvidia.com/labs/adlr/MF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07315 2025-11-11 q-fin.ST cs.AI 70%

Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation

Zonghan Wu, Congyuan Zou, Junlin Wang, Chenhan Wang, Hangjing Yang, Yilei Shao

机构 * Shanghai AI Finance School, East China Normal University(上海人工智能金融学院,华东师范大学) Tsinghua University(清华大学)

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11137 2025-11-10 cs.CL 70%

Scalable Medication Extraction and Discontinuation Identification from Electronic Health Records Using Large Language Models

Chong Shao, Douglas Snyder, Chiran Li, Bowen Gu, Kerry Ngan, Chun-Ting Yang, Jiageng Wu, Richard Wyss, Kueiyu Joshua Lin, Jie Yang

专题命中 推理评测 :reasoning(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01340 2025-11-04 cs.CV cs.CL 70%

$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles

Trishanu Das, Abhilash Nandy, Khush Bajaj, Deepiha S

机构 * Tredence Inc.(特伦德公司) Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校) Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒姆研究实验室)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments 7 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25536 2025-10-31 cs.CL 70%

TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation

Bangde Du, Minghao Guo, Songming He, Ziyi Ye, Xi Zhu, Weihang Su, Shuqi Zhu, Yujia Zhou, Yongfeng Zhang, Qingyao Ai, Yiqun Liu

机构 * Tsinghua University(清华大学) Rutgers University(罗格斯大学) Fudan University(复旦大学)

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract);分类 cs.CL

Comments Main paper: 11 pages, 3 figures, 6 tables. Appendix: 28 pages. Bangde Du and Minghao Guo contributed equally. Corresponding authors: Ziyi Ye (ziyiye@fudan.edu.cn), Qingyao Ai (aiqy@tsinghua.edu.cn)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25187 2025-10-30 cs.CL 70%

Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction

Ritesh Sunil Chavan, Jack Mostow

机构 * Department of Computer Science(计算机科学系) Stony Brook University(石溪大学) School of Computer Science(计算机科学学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23482 2025-10-28 cs.CV cs.AI 70%

On the Faithfulness of Visual Thinking: Measurement and Enhancement

Zujing Liu, Junwen Pan, Qi She, Yuan Gao, Guisong Xia

机构 * Wuhan University(武汉大学) ByteDance(字节跳动)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22455 2025-10-28 cs.SD cs.AI eess.AS 70%

Evaluating Multimodal Large Language Models on Core Music Perception Tasks

Brandon James Carone, Iran R. Roman, Pablo Ripollés

机构 * Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室) Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)

专题命中 推理评测 :reasoning(abstract);CoT(abstract);分类 cs.AI

Comments Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07458 2025-10-28 cs.CL 70%

Populism Meets AI: Advancing Populism Research with LLMs

Yujin J. Jung, Eduardo Ryô Tamaki, Julia Chatterley, Grant Mitchell, Semir Dzebo, Cristóbal Sandoval, Levente Littvay, Kirk A. Hawkins

机构 * Mount St. Mary’s University(圣玛丽大学) German Institute for Global and Area Studies(全球与区域研究所) Princeton University(普林斯顿大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Oxford(牛津大学) Diego Portales University(迪埃戈·波尔塔尔大学) ELTE Centre for Social Sciences(埃尔泰社会科学中心) Brigham Young University(Brigham Young 大学)

专题命中 推理评测 :reasoning(abstract);CoT(abstract);分类 cs.CL

Comments 27 pages, 3 figures. Preprint version under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18029 2025-10-22 cs.DB cs.AI 70%

DynaQuery: A Self-Adapting Framework for Querying Structured and Multimodal Data

Aymane Hassini

机构 * Al Akhawayn University(阿尔阿克哈文大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

Comments 15 pages, 2 figures, 10 tables. Source code and experimental artifacts are available at: https://github.com/aymanehassini/DynaQuery . The 'DynaQuery-Eval-5K' benchmark, introduced in this work, is also publicly available at: https://www.kaggle.com/datasets/aymanehassini/dynaquery-eval-5k-benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17491 2025-10-21 cs.CL 70%

Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents

Yihong Tang, Kehai Chen, Liang Yue, Jinxin Fan, Caishen Zhou, Xiaoguang Li, Yuyang Zhang, Mingming Zhao, Shixiong Kai, Kaiyang Guo, Xingshan Zeng, Wenjing Cun, Lifeng Shang, Min Zhang

机构 * School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, China(计算机科学与技术学院,哈尔滨工业大学,深圳,中国) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17052 2025-10-21 cs.AI 70%

ToolCritic: Detecting and Correcting Tool-Use Errors in Dialogue Systems

Hassan Hamad, Yingru Xu, Liang Zhao, Wenbo Yan, Narendra Gyanchandani

机构 * Amazon(亚马逊)

专题命中 推理评测 :reasoning(abstract);self-correction(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23584 2025-10-17 eess.IV cs.AI cs.CV 70%

A Clinically-Grounded Two-Stage Framework for Renal CT Report Generation

Renjie Liang, Zhengkang Fan, Jinqian Pan, Chenkun Sun, Bruce Daniel Steinberg, Russell Terry, Jie Xu

机构 * Department of Health Outcomes and Biomedical Informatics, University of Florida(健康结果与生物医学信息学系,佛罗里达大学) Department of Urology, University of Florida(泌尿外科系,佛罗里达大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04364 2025-10-16 cs.MA cs.CL 70%

Benchmarking LLMs' Swarm intelligence

Kai Ruan, Mowen Huang, Ji-Rong Wen, Hao Sun

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06134 2025-10-16 cs.CL 70%

Evaluating and Mitigating Social Bias for Large Language Models in Open-ended Settings

Zhao Liu, Tian Xie, Xueru Zhang

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12218 2025-10-15 cs.AI 70%

GOAT: A Training Framework for Goal-Oriented Agent with Tools

Hyunji Min, Sangwon Jung, Junyoung Sung, Dosung Lee, Leekyeung Han, Paul Hongsuck Seo

机构 * Korea University(韩国大学) Trillion Labs(万亿实验室)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

Comments 32 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09404 2025-10-14 cs.AI 70%

Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges

Christian Bluethgen, Dave Van Veen, Daniel Truhn, Jakob Nikolas Kather, Michael Moor, Malgorzata Polacin, Akshay Chaudhari, Thomas Frauenfelder, Curtis P. Langlotz, Michael Krauthammer, Farhad Nooralahzadeh

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04338 2025-10-07 cs.CL 70%

Evaluation of Clinical Trials Reporting Quality using Large Language Models

Mathieu Laï-king, Patrick Paroubek

机构 * Université Paris-Saclay, CNRS, LISN(巴黎-萨克雷大学、法国国家科学研究中心、LISN)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Journal ref Revue TAL 65.2, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09123 2025-10-07 cs.AI cs.CV 70%

OpenCUA: Open Foundations for Computer-Use Agents

Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Ryan Li, Xiaochuan Li, Junda Chen, Boyuan Zheng, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu, Yuyan Wang, Jixuan Chen, Yuxiao Ye, Danyang Zhang, Dikang Du, Hao Hu, Huarong Chen, Zaida Zhou, Haotian Yao, Ziwei Chen, Qizheng Gu, Yipu Wang, Heng Wang, Diyi Yang, Victor Zhong, Flood Sung, Y. Charles, Zhilin Yang, Tao Yu

机构 * XLANG Lab, The University of Hong Kong(香港大学XLANG实验室) Moonshot AI Stanford University(斯坦福大学) University of Waterloo(滑铁卢大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI

Comments Updata author list, modify first page format, correct typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03127 2025-10-06 cs.AI 70%

A Study of Rule Omission in Raven's Progressive Matrices

Binze Li

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16360 2025-10-06 cs.CL 70%

RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering

Weikang Qiu, Tinglin Huang, Ryan Rullo, Yucheng Kuang, Ali Maatouk, S. Raquel Ramos, Rex Ying

机构 * Yale University, Department of Computer Science(耶鲁大学计算机科学系) Yale University, School of Nursing(耶鲁大学护理学院) Yale University, School of Public Health(耶鲁大学公共卫生学院) Northeastern University(东北大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments ACM KDD Health Track 2025 Blue Sky Best Paper

详情

展开后加载摘要…

URL PDF HTML 收藏