arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-07 至 2025-10-07 共收录 131 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 35 篇

2510.04032 2025-10-07 cs.CL cs.AI 62%

Small Language Models for Emergency Departments Decision Support: A Benchmark Study

Zirui Wang, Jiajun Wu, Braden Teitge, Jessalyn Holodinsky, Steve Drew

机构 * Department of Electrical and Software Engineering, University of Calgary, Calgary, AB, Canada(电气与软件工程系,卡尔加里大学) Department of Emergency Medicine, University of Calgary, Calgary, AB, Canada(急诊医学系,卡尔加里大学) Rockview General Hospital, Calgary, AB, Canada(罗克维尔医院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to 2025 IEEE International Conference on Autonomous and Trusted Computing (ATC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22830 2025-10-07 cs.CL cs.AI 62%

What Has Been Lost with Synthetic Evaluation?

Alexander Gill, Abhilasha Ravichander, Ana Marasović

机构 * University of Utah(犹他大学) University of Washington(华盛顿大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments v3: Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01159 2025-10-07 cs.LG cs.AI 62%

AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science

Chenyue Li, Wen Deng, Mengqian Lu, Binhang Yuan

机构 * HKUST(香港科技大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 37 pages, 4 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18917 2025-10-07 cs.LG cs.AI 62%

Behavior Injection: Preparing Language Models for Reinforcement Learning

Zhepeng Cen, Yihang Yao, William Han, Zuxin Liu, Ding Zhao

机构 * Carnegie Mellon University(卡内基梅隆大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04934 2025-10-07 eess.AS cs.AI 57%

AURA Score: A Metric For Holistic Audio Question Answering Evaluation

Satvik Dixit, Soham Deshmukh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04623 2025-10-07 cs.AI 57%

MedPAO: A Protocol-Driven Agent for Structuring Medical Reports

Shrish Shrinath Vaidya, Gowthamaan Palani, Sidharth Ramesh, Velmurugan Balasubramanian, Minmini Selvam, Gokulraja Srinivasaraja, Ganapathy Krishnamurthi

机构 * Department of Data Science and AI, IIT Madras, India(数据科学与人工智能系,印度理工学院马德拉斯学院) Department of Engineering Design, IIT Madras, India(工程设计系,印度理工学院马德拉斯学院) LoveForm Health Technologies, India(LoveForm健康科技公司,印度) Department of Radiology and Imaging Sciences, Sri Ramachandra Institute of Higher Education and Research, India(放射学与成像科学系, Sri Ramachandra高等教育与研究学院,印度) Department of Neuro and Interventional Radiology, Sri Ramachandra Institute of Higher Education and Research, India(神经放射学与介入放射学系,Sri Ramachandra高等教育与研究学院,印度)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Paper published at "Agentic AI for Medicine" Workshop, MICCAI 2025

Journal ref Lecture Notes in Computer Science, vol 16147, 2025. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09050 2025-10-07 cs.AI 57%

ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering

Yuki Imajuku, Kohki Horie, Yoichi Iwata, Kensho Aoki, Naohiro Takahashi, Takuya Akiba

机构 * Sakana AI The University of Tokyo(东京大学) AtCoder

专题命中 推理评测 :planning(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025 Datasets & Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10610 2025-10-07 cs.CV cs.CL 57%

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman

机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) Tencent AI Seattle Lab(腾讯AI西雅图实验室) University of Edinburgh(爱丁堡大学) NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26440 2025-10-07 cs.AI 57%

Transformer Classification of Breast Lesions: The BreastDCEDL_AMBL Benchmark Dataset and 0.92 AUC Baseline

Naomi Fridman, Anat Goldstein

机构 * Department of Industrial Engineering, Ariel University(工业工程系,阿丽尔大学)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14574 2025-10-07 cs.CV cs.AI 57%

Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 推理评测 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08870 2025-10-07 cs.LG cs.MA 57%

GUIDE: Towards Scalable Advising for Research Ideas

Yaowenqi Liu, Bingxu Meng, Rui Pan, Yuxing Liu, Jerry Huang, Jiaxuan You, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18677 2025-10-07 cs.CL 57%

Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts

Eric Chamoun, Nedjma Ousidhoum, Michael Schlichtkrull, Andreas Vlachos

机构 * Department of Computer Science and Technology, University of Cambridge(计算机科学与技术系,剑桥大学) Cardiff University(卡迪夫大学) Queen Mary University of London(伦敦女王学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04261 2025-10-07 cs.CR 50%

VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy

Yu Cui, Sicheng Pan, Yifei Liu, Haibin Zhang, Cong Zuo

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03955 2025-10-07 cs.CV 50%

Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs

Sameep Vani, Shreyas Jena, Maitreya Patel, Chitta Baral, Somak Aditya, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学) Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦)

专题命中 推理评测 :reasoning(abstract)

Comments 17 pages, 9 figures, 6 tables. Presents TimeWarp, a synthetic preference data framework to improve temporal understanding in Video-LLMs, showing consistent gains across seven benchmarks. Includes supplementary material in the Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18743 2025-10-07 cs.CV 50%

SAR-TEXT: A Large-Scale SAR Image-Text Dataset Built with SAR-Narrator and A Progressive Learning Strategy for Downstream Tasks

Yiguo He, Xinjun Cheng, Junjie Zhu, Chunping Qiu, Jun Wang, Xichuan Zhang, Qiangjuan Huang, Ke Yang

机构 * Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 推理评测 :reasoning(abstract)

Comments IEEE Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08214 2025-10-07 eess.IV cs.CV 50%

Depth-Sequence Transformer (DST) for Segment-Specific ICA Calcification Mapping on Non-Contrast CT

Xiangjian Hou, Ebru Yaman Akcicek, Xin Wang, Kazem Hashemizadeh, Scott Mcnally, Chun Yuan, Xiaodong Ma

机构 * 1 Dept.\ of Electrical \& Computer Engineering, University of Utah, Salt Lake City, UT, USA 2 Dept.\ of Radiology \& Imaging Sciences, University of Utah, Salt Lake City, UT, USA 3 Dept.\ of Electrical \& Computer Engineering, University of Washington, Seattle, WA, USA

专题命中 推理评测 :planning(abstract)

Comments Accept to IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15778 2025-10-07 cs.CV cs.RO 50%

AutoDrive-QA: A Multiple-Choice Benchmark for Vision-Language Evaluation in Urban Autonomous Driving

Boshra Khalili, Andrew W. Smyth

机构 * Columbia University(哥伦比亚大学)

专题命中 推理评测 :planning(abstract)

Comments Updated results with larger dataset experiments and expanded discussion

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 24 篇

2510.04045 2025-10-07 cs.CL cs.LG 90%

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

Yunfan Zhang, Kathleen McKeown, Smaranda Muresan

机构 * Columbia University(哥伦比亚大学) Barnard College(巴纳德学院)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.LG

Comments ACL EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04871 2025-10-07 cs.LG cs.AI 81%

Less is More: Recursive Reasoning with Tiny Networks

Alexia Jolicoeur-Martineau

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03264 2025-10-07 cs.LG cs.AI 81%

Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data

Syeda Nahida Akter, Shrimai Prabhumoye, Eric Nyberg, Mostofa Patwary, Mohammad Shoeybi, Yejin Choi, Bryan Catanzaro

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学) Stanford University(斯坦福大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03598 2025-10-07 cs.CV cs.LG 79%

Exploring the Hierarchical Reasoning Model for Small Natural-Image Classification Without Augmentation

Alexander V. Mantzaris

机构 * Department of Data, Mathematical and Statistical Sciences, University of Central Florida(数据、数学与统计科学系,中央佛罗里达大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02608 2025-10-07 cs.AI 79%

Mitigating Modal Imbalance in Multimodal Reasoning

Chen Henry Wu, Neil Kale, Aditi Raghunathan

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI

Comments 10 pages, 10 figures, CoLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09532 2025-10-07 cs.RO cs.AI 79%

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

Congcong Wen, Geeta Chandra Raju Bethala, Yu Hao, Niraj Pudasaini, Hao Huang, Shuaihang Yuan, Baoru Huang, Anh Nguyen, Mengyu Wang, Anthony Tzes, Yi Fang

机构 * Embodied AI and Robotics (AIR) Lab, New York University, New York, USA and NYUAD Center for Artificial Intelligence and Robotics, New York University Abu Dhabi, Abu Dhabi, UAE(纽约大学Embodied AI和机器人实验室及纽约大学阿布扎赫尔人工智能与机器人中心) Harvard AI and Robotics Lab, Harvard University, Boston, USA(哈佛大学人工智能与机器人实验室) Department of Computer Science, University College London, London, UK(伦敦大学学院计算机科学系) Department of Computer Science, University of Liverpool, UK(利物浦大学计算机科学系) NYUAD Center for Artificial Intelligence and Robotics, New York University Abu Dhabi, Abu Dhabi, UAE(纽约大学阿布扎赫尔人工智能与机器人中心)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI

Comments website link: https://humanoid-coa.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04454 2025-10-07 cs.CL 77%

Mitigating Forgetting Between Supervised and Reinforcement Learning Yields Stronger Reasoners

Xiangchi Yuan, Xiang Chen, Tong Yu, Dachuan Shi, Can Jin, Wenke Lee, Saayan Mitra

机构 * Georgia Institute of Technology(佐治亚理工学院) Adobe Research(Adobe研究) Rutgers University(罗格斯大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04488 2025-10-07 cs.AI cs.IT math.IT 74%

Multi-Agent Collaborative Intelligence: Dual-Dial Control for Reliable LLM Reasoning

Edward Y. Chang, Ethan Y. Chang

机构 * Stanford University(斯坦福大学) UIUC(伊利诺伊大学香槟分校)

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments 27 pages, 5 figures, 21 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16856 2025-10-07 cs.CV cs.AI 70%

SIA: Enhancing Safety via Intent Awareness for Vision-Language Models

Youngjin Na, Sangheon Jeong, Youngwan Lee, Jian Lee, Dawoon Jeong, Youngman Kim

机构 * VLM Safety LAB, MODULABS(视觉语言模型安全实验室,MODULABS) ETRI(电子技术研究院) KAIST(韩国科学技术院)

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI

Comments Accepted to Safe and Trustworthy Multimodal AI Systems(SafeMM-AI) Workshop at ICCV2025, Non-archival track

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20783 2025-10-07 cs.LG cs.AI cs.CL 67%

Understanding R1-Zero-Like Training: A Critical Perspective

Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin

机构 * Sea AI Lab(海智实验室) National University of Singapore(新加坡国立大学) Singapore Management University(新加坡管理学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04311 2025-10-07 cs.AI cs.LG 62%

On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems

Bohan Tang, Huidong Liang, Keyue Jiang, Xiaowen Dong

机构 * University of Oxford(牛津大学) University College London(伦敦大学学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04196 2025-10-07 cs.AI cs.LG 62%

COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability

Yizhuo Ding, Mingkang Chen, Qiuhua Liu, Fenghua Weng, Wanying Qu, Yue Yang, Yugang Jiang, Zuxuan Wu, Yanwei Fu, Wenqi Shao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) ShanghaiTech University(上海交通大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03521 2025-10-07 cs.CL cs.AI 62%

Identifying Financial Risk Information Using RAG with a Contrastive Insight

Ali Elahi

机构 * Department of Computer Science University of Illinois Chicago(伊利诺伊大学芝加哥分校计算机科学系) Surlamer Investments(Surlamer投资公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 1 figure, Workshop on Generative AI in Finance, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏