arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-06 至 2025-10-06 共收录 15 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 15 篇

2510.02351 2025-10-06 cs.CL cs.AI 81%

Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs

Dzmitry Pihulski, Jan Kocoń

机构 * Department of Artificial Intelligence, Wroclaw Tech, Poland(人工智能系,沃拉布技术学院,波兰)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments To appear in the Proceedings of the IEEE International Conference on Data Mining Workshops (ICDMW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02892 2025-10-06 cs.LG 80%

RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning

Aleksei Arzhantsev, Otmane Sakhi, Flavian Vasile

机构 * Criteo AI Lab(Criteo人工智能实验室) Ecole Polytechnique Paris(巴黎高等理工学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

Comments Accepted to the Efficient Reasoning Workshop at NeuRIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02780 2025-10-06 cs.CV 80%

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

Prahitha Movva

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 推理评测 :reasoning(title,abstract);planning(journal_ref)

Journal ref COLM 2025: First Workshop on the Application of LLM Explainability to Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02906 2025-10-06 q-fin.CP cs.AI 79%

FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence

Abhinav Arun, Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali

机构 * Domyn

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00579 2025-10-06 cs.MM cs.IR 78%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 推理评测 :reasoning(title,abstract)

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03127 2025-10-06 cs.AI 70%

A Study of Rule Omission in Raven's Progressive Matrices

Binze Li

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16360 2025-10-06 cs.CL 70%

RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering

Weikang Qiu, Tinglin Huang, Ryan Rullo, Yucheng Kuang, Ali Maatouk, S. Raquel Ramos, Rex Ying

机构 * Yale University, Department of Computer Science(耶鲁大学计算机科学系) Yale University, School of Nursing(耶鲁大学护理学院) Yale University, School of Public Health(耶鲁大学公共卫生学院) Northeastern University(东北大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments ACM KDD Health Track 2025 Blue Sky Best Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02830 2025-10-06 cs.CL cs.AI 62%

Evaluating Large Language Models for IUCN Red List Species Information

Shinya Uryu

机构 * Center for Design-Oriented AI Education and Research(设计导向人工智能教育与研究中心) Tokushima University(德岛大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02549 2025-10-06 cs.CL cs.AI 62%

Knowledge-Graph Based RAG System Evaluation Framework

Sicheng Dong, Vahid Zolfaghari, Nenad Petrovic, Alois Knoll

机构 * Technical University of Munich, Robotics, Artificial Intelligence and Embedded Systems(慕尼黑技术大学,机器人,人工智能与嵌入式系统)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02326 2025-10-06 cs.CL cs.AI 62%

Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval

Vivek Bhavsar, Joseph Ereifej, Aravanan Gurusami

机构 * CTO Office, Coherent Corporation(Coherent Corporation 技术总监办公室)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01227 2025-10-06 cs.CL cs.LG math.HO 62%

EEFSUVA: A New Mathematical Olympiad Benchmark

Nicole N Khatibi, Daniil A. Radamovich, Michael P. Brenner

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

Comments 16 Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14052 2025-10-06 cs.IR cs.AI cs.CL 62%

FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering

Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira, Chaewoon Kim, Minjae Kim, Juneha Hwang, Jaeseon Ha, Hojun Choi, Suyeol Yun, Yongjin Kim, Yongjae Lee

机构 * University of Florida(佛罗里达大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02325 2025-10-06 cs.CR cs.AI 57%

Agentic-AI Healthcare: Multilingual, Privacy-First Framework with MCP Agents

Mohammed A. Shehab

机构 * Concordia Continuing Education(康科德继续教育) Concordia University(康科德大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 1 figure. Submitted as a system/vision paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09650 2025-10-06 cs.CV cs.LG cs.MM cs.RO eess.IV 57%

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen, Junwei Zheng, Yufan Chen, Kailun Yang, Jiamin Wu, Chongqing Hao, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) HEBUST

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02534 2025-10-06 cs.SE 50%

ZeroFalse: Improving Precision in Static Analysis with LLMs

Mohsen Iranmanesh, Sina Moradi Sabet, Sina Marefat, Ali Javidi Ghasr, Allison Wilson, Iman Sharafaldin, Mohammad A. Tayebi

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏