arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10549 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10549 篇

2501.19017 2025-10-09 cs.CL 57%

Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models

Bin Zhu, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学) Fudan University(复旦大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Project website: https://yxg1005.github.io/GaslightingNegationAttacks/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05458 2025-10-08 cs.CL 57%

SocialNLI: A Dialogue-Centric Social Inference Dataset

Akhil Deo, Kate Sanders, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05362 2025-10-08 cs.CL 57%

Residualized Similarity for Faithfully Explainable Authorship Verification

Peter Zeng, Pegah Alipoormolabashi, Jihu Mun, Gourab Dey, Nikita Soni, Niranjan Balasubramanian, Owen Rambow, H. Schwartz

机构 * Department of Computer Science(计算机科学系) Department of Linguistics(语言学系) Institute for Advanced Computational Science(先进计算科学研究院) Stony Brook University(石溪大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05046 2025-10-08 cs.CL 57%

COLE: a Comprehensive Benchmark for French Language Understanding Evaluation

David Beauchemin, Yan Tremblay, Mohamed Amine Youssef, Richard Khoury

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Submitted to ACL Rolling Review of October

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09098 2025-10-08 cs.CL 57%

SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models

Kehua Feng, Xinyi Shen, Weijie Wang, Xiang Zhuang, Yuqi Tang, Qiang Zhang, Keyan Ding

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University(浙江大学Hangzhou全球科技创新中心) ZJU-UIUC Institute, Zhejiang University(浙江大学UIUC研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 33 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04934 2025-10-07 eess.AS cs.AI 57%

AURA Score: A Metric For Holistic Audio Question Answering Evaluation

Satvik Dixit, Soham Deshmukh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04623 2025-10-07 cs.AI 57%

MedPAO: A Protocol-Driven Agent for Structuring Medical Reports

Shrish Shrinath Vaidya, Gowthamaan Palani, Sidharth Ramesh, Velmurugan Balasubramanian, Minmini Selvam, Gokulraja Srinivasaraja, Ganapathy Krishnamurthi

机构 * Department of Data Science and AI, IIT Madras, India(数据科学与人工智能系,印度理工学院马德拉斯学院) Department of Engineering Design, IIT Madras, India(工程设计系,印度理工学院马德拉斯学院) LoveForm Health Technologies, India(LoveForm健康科技公司,印度) Department of Radiology and Imaging Sciences, Sri Ramachandra Institute of Higher Education and Research, India(放射学与成像科学系, Sri Ramachandra高等教育与研究学院,印度) Department of Neuro and Interventional Radiology, Sri Ramachandra Institute of Higher Education and Research, India(神经放射学与介入放射学系,Sri Ramachandra高等教育与研究学院,印度)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Paper published at "Agentic AI for Medicine" Workshop, MICCAI 2025

Journal ref Lecture Notes in Computer Science, vol 16147, 2025. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09050 2025-10-07 cs.AI 57%

ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering

Yuki Imajuku, Kohki Horie, Yoichi Iwata, Kensho Aoki, Naohiro Takahashi, Takuya Akiba

机构 * Sakana AI The University of Tokyo(东京大学) AtCoder

专题命中 推理评测 :planning(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025 Datasets & Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10610 2025-10-07 cs.CV cs.CL 57%

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman

机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) Tencent AI Seattle Lab(腾讯AI西雅图实验室) University of Edinburgh(爱丁堡大学) NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26440 2025-10-07 cs.AI 57%

Transformer Classification of Breast Lesions: The BreastDCEDL_AMBL Benchmark Dataset and 0.92 AUC Baseline

Naomi Fridman, Anat Goldstein

机构 * Department of Industrial Engineering, Ariel University(工业工程系,阿丽尔大学)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14574 2025-10-07 cs.CV cs.AI 57%

Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08870 2025-10-07 cs.LG cs.MA 57%

GUIDE: Towards Scalable Advising for Research Ideas

Yaowenqi Liu, Bingxu Meng, Rui Pan, Yuxing Liu, Jerry Huang, Jiaxuan You, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18677 2025-10-07 cs.CL 57%

Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts

Eric Chamoun, Nedjma Ousidhoum, Michael Schlichtkrull, Andreas Vlachos

机构 * Department of Computer Science and Technology, University of Cambridge(计算机科学与技术系,剑桥大学) Cardiff University(卡迪夫大学) Queen Mary University of London(伦敦女王学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02325 2025-10-06 cs.CR cs.AI 57%

Agentic-AI Healthcare: Multilingual, Privacy-First Framework with MCP Agents

Mohammed A. Shehab

机构 * Concordia Continuing Education(康科德继续教育) Concordia University(康科德大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 1 figure. Submitted as a system/vision paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09650 2025-10-06 cs.CV cs.LG cs.MM cs.RO eess.IV 57%

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen, Junwei Zheng, Yufan Chen, Kailun Yang, Jiamin Wu, Chongqing Hao, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) HEBUST

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01887 2025-10-03 q-fin.CP cs.AI 57%

FINCH: Financial Intelligence using Natural language for Contextualized SQL Handling

Avinash Kumar Singh, Bhaskarjit Sarmah, Stefano Pasquali

机构 * Domyn

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01558 2025-10-03 cs.CE cs.LG eess.SP 57%

CardioRAG: A Retrieval-Augmented Generation Framework for Multimodal Chagas Disease Detection

Zhengyang Shen, Xuehao Zhai, Hua Tu, Mayue Shi

机构 * Department of Electrical and Electronic Engineering Imperial College London(帝国理工学院电子与电气工程系) Department of Civil and Environmental Engineering Imperial College London(帝国理工学院土木与环境工程系) Institute of Biomedical Engineering Department of Engineering Science University of Oxford(牛津大学生物医学工程研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 4 pages, 2 figures. Accepted for oral presentation at the 52nd international Computing in Cardiology Conference (CinC2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01398 2025-10-03 cs.AI 57%

Automating Data-Driven Modeling and Analysis for Engineering Applications using Large Language Model Agents

Yang Liu, Zaid Abulawi, Abhiram Garimidi, Doyeong Lim

机构 * Department of Nuclear Engineering, Texas A\&M University(核工程系,德克萨斯A&M大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01249 2025-10-03 cs.CL 57%

LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning

You-Le Fang, Dong-Shan Jian, Xiang Li, Ce Meng, Ling-Shi Meng, Chen-Xu Yan, Zhi-Zhang Bian, Yan-Qing Ma

机构 * School of Physics, Peking University(北京大学物理学院) Center for High Energy Physics, Peking University(北京大学高等能源物理中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 29 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00311 2025-10-02 cs.CL 57%

CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage

Bowen Wei, Yuan Shen Tay, Howard Liu, Jinhao Pan, Kun Luo, Ziwei Zhu, Chris Jordan

机构 * George Mason University(乔治·马歇尔大学) Fluency Security(流畅安全)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00161 2025-10-02 cs.CL 57%

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 21 pages. Code: https://github.com/kimihiroh/tama

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26225 2025-10-01 cs.CV cs.AI 57%

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris

机构 * IEEE CBMI 2025

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25229 2025-10-01 cs.AI 57%

Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models

Lukas Petersson, Axel Backlund, Axel Wennstöm, Hanna Petersson, Callum Sharrock, Arash Dabiri

机构 * Andon Labs(安登实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 9 pages, 8 figures, submitted for ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04909 2025-10-01 cs.CV cs.AI 57%

HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding

Yuxuan Cai, Jiangning Zhang, Zhenye Gan, Qingdong He, Xiaobin Hu, Junwei Zhu, Yabiao Wang, Chengjie Wang, Zhucun Xue, Chaoyou Fu, Xinwei He, Xiang Bai

机构 * Fantasyele

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24350 2025-09-30 cs.CV cs.AI 57%

Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA

Yan Ke, Xin Yu, Heming Du, Scott Chapman, Helen Huang

机构 * The University of Queensland(昆士兰大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 13 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20707 2025-09-30 cs.AI 57%

An Automated Retrieval-Augmented Generation LLaMA-4 109B-based System for Evaluating Radiotherapy Treatment Plans

Junjie Cui, Peilong Wang, Jason Holmes, Leshan Sun, Michael L. Hinni, Barbara A. Pockaj, Sujay A. Vora, Terence T. Sio, William W. Wong, Nathan Y. Yu, Steven E. Schild, Joshua R. Niska, Sameer R. Keole, Jean-Claude M. Rwigema, Samir H. Patel, Lisa A. McGee, Carlos A. Vargas, Wei Liu

机构 * Mayo Clinic Arizona(梅奥诊所亚利桑那分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 16 pages, 4 figures. Submitted to npj Digital Medicine

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24217 2025-09-30 cs.CL 57%

Semi-structured LLM Reasoners Can Be Rigorously Audited

Jixuan Leng, Cassandra A. Cohen, Zhixian Zhang, Chenyan Xiong, William W. Cohen

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17127 2025-09-30 cs.CV cs.LG 57%

Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts

Michal Golovanevsky, William Rudman, Michael Lepori, Amir Bar, Ritambhara Singh, Carsten Eickhoff

机构 * Brown University(布朗大学) Tel Aviv University(特拉维夫大学) University of Tübingen(图宾根大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04740 2025-09-30 cs.CV cs.AI 57%

SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data

Samarth Mishra, Kate Saenko, Venkatesh Saligrama

机构 * Boston University(波士顿大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23827 2025-09-30 cs.CV cs.LG 57%

Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models

Efthymios Tsaprazlis, Tiantian Feng, Anil Ramakrishna, Rahul Gupta, Shrikanth Narayanan

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏