arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-08 至 2025-10-08 共收录 23 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 23 篇

2510.05871 2025-10-08 cs.AI cs.LG 90%

Towards Label-Free Biological Reasoning Synthetic Dataset Creation via Uncertainty Filtering

Josefa Lia Stoisser, Lawrence Phillips, Aditya Misra, Tom A. Lamb, Philip Torr, Marc Boubnovski Martell, Julien Fauqueur, Kaspar Märtens

机构 * Novo Nordisk(诺华制药) University of Oxford(牛津大学)

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title);chain-of-thought(abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06217 2025-10-08 cs.AI cs.CL cs.LG 82%

TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning

Jiaru Zou, Soumya Roy, Vinay Kumar Verma, Ziyi Wang, David Wipf, Pan Lu, Sumit Negi, James Zou, Jingrui He

机构 * UIUC(伊利诺伊大学香槟分校) Amazon(亚马逊) Purdue University(普渡大学) Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05547 2025-10-08 cs.RO 82%

ARRC: Advanced Reasoning Robot Control - Knowledge-Driven Autonomous Manipulation Using Retrieval-Augmented Generation

Eugene Vorobiov, Ammar Jaleel Mahmood, Salim Rezvani, Robin Chhabra

机构 * Department of Mechanical, Industrial and Mechatronics Engineering, Toronto Metropolitan University(机械、工业与机电工程系,多伦多 Metropolitan 大学)

专题命中 推理评测 :reasoning(title,abstract);planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07086 2025-10-08 cs.LG cs.CL 81%

A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility

Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, Matthias Bethge

机构 * Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) University of Cambridge(剑桥大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06014 2025-10-08 cs.AI 79%

ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models

Zhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Zhiyuan Yu, Qipeng Guo, Xuanjing Huang, Xipeng Qiu

机构 * Fudan University(复旦大学) The University of Hong Kong(香港大学) Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05335 2025-10-08 cs.AI 79%

Biomedical reasoning in action: Multi-agent System for Auditable Biomedical Evidence Synthesis

Oskar Wysocki, Magdalena Wysocka, Mauricio Jacobo, Harriet Unsworth, André Freitas

机构 * Idiap Research Institute(IDiap研究 institute) National Biomarker Centre (NBC) CRUK Manchester Institute(国家生物标记中心(NBC)CRUK曼彻斯特研究所) Department of Computer Science University of Manchester, UK(计算机科学系曼彻斯特大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22584 2025-10-08 cs.LG cs.AI cs.CL 75%

BenchAgents: Multi-Agent Systems for Structured Benchmark Creation

Natasha Butt, Varun Chandrasekaran, Neel Joshi, Besmira Nushi, Vidhisha Balachandran

机构 * University of Amsterdam(阿姆斯特丹大学) Microsoft Research(微软研究院) UIUC(伊利诺伊大学香槟分校)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02563 2025-10-08 cs.LG cs.CL 73%

DynaGuard: A Dynamic Guardian Model With User-Defined Policies

Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph Vincent, Chirag Jain, Melissa Kazemi Rad, C. Bayan Bruss, Ashwinee Panda, Tom Goldstein

机构 * University of Maryland(马里兰大学) Capital One

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

Comments 22 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05935 2025-10-08 cs.LG cs.AI 62%

LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection

Mohamed Bal-Ghaoui, Fayssal Sabri

机构 * R&D Department, Audensiel Conseil(Audensiel Conseil 研发部) Ecole Centrale de Lyon(Lyon 工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09387 2025-10-08 cs.LG cs.AI 62%

MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization

Mohamed Bal-Ghaoui, Mohammed Tiouti

机构 * R&D Department, Audensiel Conseil Paris, France(Audensiel Conseil 巴黎法国研发部) Université Évry Paris-Saclay(Évry 巴黎-萨克莱大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21432 2025-10-08 cs.CL cs.AI 62%

Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour

Tareq Alsaleh, Bilal Farooq

机构 * Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University, Canada(交通创新实验室(LiTrans)、多伦多 Metropolitan 大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05151 2025-10-08 cs.CL cs.LG 62%

Exploring Large Language Models for Financial Applications: Techniques, Performance, and Challenges with FinMA

Prudence Djagba, Abdelkader Y. Saley

机构 * Lyman Briggs College, Michigan State University(密歇根州立大学Lyman Briggs学院) Department of Finance, Michigan State University(密歇根州立大学金融系) African Institute for Mathematical Sciences, Rwanda(刚果(金)数学科学研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05135 2025-10-08 cs.CL cs.LG 62%

Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment

Vanya Bannihatti Kumar, Divyanshu Goyal, Akhil Eppa, Neel Bhandari

机构 * Adobe Inc.(Adobe公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22385 2025-10-08 cs.CV cs.AI cs.CL 62%

Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment

Yue Zhang, Jilei Sun, Yunhui Guo, Vibhav Gogate

机构 * Department of Computer Science(计算机科学系) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05458 2025-10-08 cs.CL 57%

SocialNLI: A Dialogue-Centric Social Inference Dataset

Akhil Deo, Kate Sanders, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05362 2025-10-08 cs.CL 57%

Residualized Similarity for Faithfully Explainable Authorship Verification

Peter Zeng, Pegah Alipoormolabashi, Jihu Mun, Gourab Dey, Nikita Soni, Niranjan Balasubramanian, Owen Rambow, H. Schwartz

机构 * Department of Computer Science(计算机科学系) Department of Linguistics(语言学系) Institute for Advanced Computational Science(先进计算科学研究院) Stony Brook University(石溪大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05046 2025-10-08 cs.CL 57%

COLE: a Comprehensive Benchmark for French Language Understanding Evaluation

David Beauchemin, Yan Tremblay, Mohamed Amine Youssef, Richard Khoury

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Submitted to ACL Rolling Review of October

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09098 2025-10-08 cs.CL 57%

SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models

Kehua Feng, Xinyi Shen, Weijie Wang, Xiang Zhuang, Yuqi Tang, Qiang Zhang, Keyan Ding

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University(浙江大学Hangzhou全球科技创新中心) ZJU-UIUC Institute, Zhejiang University(浙江大学UIUC研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 33 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05533 2025-10-08 q-fin.PM 50%

The New Quant: A Survey of Large Language Models in Financial Prediction and Trading

Weilong Fu

专题命中 推理评测 :reasoning(abstract)

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05365 2025-10-08 cs.SE 50%

Test Case Generation from Bug Reports via Large Language Models: A Cognitive Layered Evaluation Framework

Irtaza Sajid Qureshi, Zhen Ming, Jiang

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08833 2025-10-08 cs.CY 50%

Position: The Pitfalls of Over-Alignment: Overly Caution Health-Related Responses From LLMs are Unethical and Dangerous

Wenqi Marshall Guo, Yiyang Du, Heidi J. S. Tworek, Shan Du

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19271 2025-10-08 cs.SE 50%

VerifyThisBench: Generating Code, Specifications, and Proofs All at Once

Xun Deng, Sicheng Zhong, Barış Bayazıt, Andreas Veneris, Fan Long, Xujie Si

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15192 2025-10-08 cs.CV 50%

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

Fatemeh Ziaeetabar, Florentin Wörgötter

机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,塔里斯坦大学)

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏