arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-07-30 至 2025-07-30 共收录 46 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 16 篇

2507.08128 2025-07-30 cs.SD cs.AI cs.CL eess.AS 73%

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Arushi Goel, Sreyan Ghosh, Jaehyeon Kim, Sonal Kumar, Zhifeng Kong, Sang-gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, Bryan Catanzaro

机构 * NVIDIA, USA(NVIDIA公司) University of Maryland, College Park, USA(马里兰大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments Code, Datasets, and Models: https://research.nvidia.com/labs/adlr/AF3/ ; Updates in v2: Updated results for new thinking mode ckpts, added qualitative figure, added note on fully open claim, add email ID for corresponding authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21585 2025-07-30 cs.AI 70%

SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation

Hao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu, Huadong Ma

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22034 2025-07-30 cs.AI cs.CL cs.LG 67%

UserBench: An Interactive Gym Environment for User-Centric Agents

Cheng Qian, Zuxin Liu, Akshara Prabhakar, Zhiwei Liu, Jianguo Zhang, Haolin Chen, Heng Ji, Weiran Yao, Shelby Heinecke, Silvio Savarese, Caiming Xiong, Huan Wang

机构 * Salesforce AI Research(Salesforce AI研究部) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 Pages, 17 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21188 2025-07-30 cs.LG cs.AI 62%

Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs

Raj Krishnan Vijayaraj

机构 * Independent Researcher(独立研究者)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21828 2025-07-30 cs.CL 57%

Modelling Adjectival Modification Effects on Semantic Plausibility

Anna Golub, Beate Zywietz, Annerose Eichel

机构 * Institute for Natural Language Processing, University of Stuttgart(自然语言处理研究所,斯图加特大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted at ESSLLI 2025 Student Session

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21428 2025-07-30 cs.CL 57%

MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations

Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, James A. Burke

机构 * Commercial Technology and Innovation Office, PricewaterhouseCoopers(普华永道商业技术与创新办公室)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 23 Pages, 20 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21287 2025-07-30 cs.AI 57%

Structured Relevance Assessment for Robust Retrieval-Augmented Language Models

Aryan Raj, Astitva Veer Garg, Anitha D

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments International Conference on ICT for Sustainable Development (ICT4SD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21130 2025-07-30 cs.AI 57%

INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems

Bintao Tang, Xin Yang, Yuhao Wang, Zixuan Qiu, Zimo Ji, Wenyuan Jiang

机构 * School of Software Engineering, Tongji University, Shanghai, China(同济大学软件工程学院) Polytechnic Institute, Zhejiang University, Zhejiang, China(浙江大学 polytechnic 院) ETH Zurich, Zurich, Switzerland(苏黎世联邦理工学院) Department of Computer Science(计算机科学系) Engineering, Hong Kong University of Science(科学大学工程学院) School of Mathematics(数学学院) Physics, Xi’an Jiaotong Liverpool University, Suzhou, China(西安交通大学利物浦大学物理学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 19 pages, 5 figures

Journal ref 2nd AI for Math Workshop @ ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16491 2025-07-30 cs.CL 57%

BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data

Wenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou, Mona Diab, Maarten Sap

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14540 2025-07-30 cs.RO cs.AI cs.CV 57%

IRASim: A Fine-Grained World Model for Robot Manipulation

Fangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu, Chilam Cheang, Tao Kong

机构 * Hong Kong University of Science and Technology(香港科技大学) ByteDance Seed(字节跳动种子)

专题命中 推理评测 :planning(abstract);分类 cs.AI

Comments Opensource, project website: https://gen-irasim.github.io

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21969 2025-07-30 cs.MA 50%

Towards Cognitive Synergy in LLM-Based Multi-Agent Systems: Integrating Theory of Mind and Critical Evaluation

Adam Kostka, Jarosław A. Chudziak

专题命中 推理评测 :reasoning(abstract)

Comments Accepted at CogSci 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 50%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21113 2025-07-30 cs.CR 50%

Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Farzana Abdulzada

专题命中 推理评测 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 3 篇

2506.05413 2025-07-30 cs.CL cs.AI cs.LG 67%

SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs

Patrik Czakó, Gábor Kertész, Sándor Szénási

机构 * Doctoral School of Applied Informatics and Applied Mathematics, Obuda University(应用信息学与应用数学博士学院,奥布达大学) John von Neumann Faculty of Informatics, Obuda University(约翰·冯·诺依曼信息学院,奥布达大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 6 pages, 3 figures, 5 tables. Accepted to IEEE SMC 2025 conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00378 2025-07-30 cs.SE cs.AI 57%

iPanda: An LLM-based Agent for Automated Conformance Testing of Communication Protocols

Xikai Sun, Fan Dang, Shiqi Jiang, Jingao Xu, Kebin Liu, Xin Miao, Zihao Yang, Weichen Zhang, Haimo Lu, Yawen Zheng, Yunhao Liu

机构 * Tsinghua University(清华大学) Microsoft Research Asia(微软亚洲研究院) Carnegie Mellon University(卡内基梅隆大学) Yanshan University(燕山大学)

专题命中 其他推理 :CoT(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18985 2025-07-30 cs.CV cs.AI 57%

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Guanxi Shen

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

Comments Keywords: Explainable Computer Vision, Large Vision-Language Models, AI Interpretability, Explainable AI, Visual Saliency, Attribution Maps, Cross-Modal Attribution, Human Attention Alignment, AI Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏