arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-06 至 2025-10-06 共收录 59 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 6 篇

2510.02850 2025-10-06 cs.AI 57%

Reward Model Routing in Alignment

Xinle Wu, Yao Lu

机构 * National University of Singapore(新加坡国立大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 复杂问题求解 9 篇

2510.02827 2025-10-06 cs.CL cs.IR 83%

StepChain GraphRAG: Reasoning Over Knowledge Graphs for Multi-Hop Question Answering

Tengjun Ni, Xin Yuan, Shenghong Li, Kai Wu, Ren Ping Liu, Wei Ni, Wenjie Zhang

机构 * University of Technology Sydney, Australia(澳大利亚技术大学) Data61, CSIRO, Australia(数据61,CSIRO) School of Engineering, Edith Cowan University, Australia(埃迪斯·科温大学工程学院) University of New South Wales, Australia(新南威尔士大学)

专题命中 复杂问题求解 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02648 2025-10-06 cs.CL 79%

SoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language Models

Rui Qi, Zhibo Man, Yufeng Chen, Fengran Mo, Jinan Xu, Kaiyu Huang

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通 key laboratory(北京交通大学)) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院(北京交通大学)) University of Montreal(蒙特利尔大学)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL

Comments EMNLP 2025 (findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00496 2025-10-06 cs.CL 79%

Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations

Pengzhou Cheng, Lingzhong Dong, Zeng Wu, Zongru Wu, Xiangru Tang, Chengwei Qin, Zhuosheng Zhang, Gongshen Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Yale University(耶鲁大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL

Comments 23 pages, 10 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02557 2025-10-06 cs.AI 70%

Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge

Charlie Masters, Advaith Vellanki, Jiangbo Shangguan, Bart Kultys, Jonathan Gilmore, Alastair Moore, Stefano V. Albrecht

机构 * DeepFlow(深流)

专题命中 复杂问题求解 :reasoning(abstract);planning(abstract);分类 cs.AI

Comments Accepted as an oral paper for the conference for Distributed Artificial Intelligence (DAI 2025). 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02719 2025-10-06 cs.CL cs.AI 62%

TravelBench : Exploring LLM Performance in Low-Resource Domains

Srinivas Billa, Xiaonan Jing

机构 * Expedia Group(Expedia集团)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05522 2025-10-06 cs.LG cs.AI 62%

Continuous Thought Machines

Luke Darlow, Ciaran Regan, Sebastian Risi, Jeffrey Seely, Llion Jones

机构 * Sakana AI University of Tsukuba(东京大学) IT University of Copenhagen(哥本哈根IT大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Technical report accompanied by online project page: https://pub.sakana.ai/ctm/ Accepted as a spotlight paper at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11191 2025-10-06 cs.CR cs.AI cs.CL 62%

Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training

Yao-Ching Yu, Tsun-Han Chiang, Cheng-Wei Tsai, Chien-Ming Huang, Wen-Kwang Tsao

机构 * AI Lab, TrendMicro(TrendMicro人工智能实验室)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23899 2025-10-06 cs.CV 50%

Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq, Rakesh Chandra Joshi

机构 * Amity University(阿米蒂大学)

专题命中 复杂问题求解 :reasoning(abstract)

Comments 12 pages (9 main + 2 references/appendix), 2 figures, conference paper submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20088 2025-10-06 cs.CV cs.MM cs.SD 50%

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma, Yixiao Ge, Wei Zou, Ying Shan

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) MAIS, Institute of Automation, CAS, Beijing(自动化研究所北京研究所MAIS)

专题命中 复杂问题求解 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 推理评测 15 篇

2510.02351 2025-10-06 cs.CL cs.AI 81%

Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs

Dzmitry Pihulski, Jan Kocoń

机构 * Department of Artificial Intelligence, Wroclaw Tech, Poland(人工智能系,沃拉布技术学院,波兰)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments To appear in the Proceedings of the IEEE International Conference on Data Mining Workshops (ICDMW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02892 2025-10-06 cs.LG 80%

RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning

Aleksei Arzhantsev, Otmane Sakhi, Flavian Vasile

机构 * Criteo AI Lab(Criteo人工智能实验室) Ecole Polytechnique Paris(巴黎高等理工学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

Comments Accepted to the Efficient Reasoning Workshop at NeuRIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02780 2025-10-06 cs.CV 80%

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

Prahitha Movva

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 推理评测 :reasoning(title,abstract);planning(journal_ref)

Journal ref COLM 2025: First Workshop on the Application of LLM Explainability to Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02906 2025-10-06 q-fin.CP cs.AI 79%

FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence

Abhinav Arun, Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali

机构 * Domyn

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00579 2025-10-06 cs.MM cs.IR 78%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 推理评测 :reasoning(title,abstract)

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03127 2025-10-06 cs.AI 70%

A Study of Rule Omission in Raven's Progressive Matrices

Binze Li

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16360 2025-10-06 cs.CL 70%

RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering

Weikang Qiu, Tinglin Huang, Ryan Rullo, Yucheng Kuang, Ali Maatouk, S. Raquel Ramos, Rex Ying

机构 * Yale University, Department of Computer Science(耶鲁大学计算机科学系) Yale University, School of Nursing(耶鲁大学护理学院) Yale University, School of Public Health(耶鲁大学公共卫生学院) Northeastern University(东北大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments ACM KDD Health Track 2025 Blue Sky Best Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02830 2025-10-06 cs.CL cs.AI 62%

Evaluating Large Language Models for IUCN Red List Species Information

Shinya Uryu

机构 * Center for Design-Oriented AI Education and Research(设计导向人工智能教育与研究中心) Tokushima University(德岛大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02549 2025-10-06 cs.CL cs.AI 62%

Knowledge-Graph Based RAG System Evaluation Framework

Sicheng Dong, Vahid Zolfaghari, Nenad Petrovic, Alois Knoll

机构 * Technical University of Munich, Robotics, Artificial Intelligence and Embedded Systems(慕尼黑技术大学,机器人,人工智能与嵌入式系统)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02326 2025-10-06 cs.CL cs.AI 62%

Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval

Vivek Bhavsar, Joseph Ereifej, Aravanan Gurusami

机构 * CTO Office, Coherent Corporation(Coherent Corporation 技术总监办公室)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01227 2025-10-06 cs.CL cs.LG math.HO 62%

EEFSUVA: A New Mathematical Olympiad Benchmark

Nicole N Khatibi, Daniil A. Radamovich, Michael P. Brenner

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

Comments 16 Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14052 2025-10-06 cs.IR cs.AI cs.CL 62%

FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering

Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira, Chaewoon Kim, Minjae Kim, Juneha Hwang, Jaeseon Ha, Hojun Choi, Suyeol Yun, Yongjin Kim, Yongjae Lee

机构 * University of Florida(佛罗里达大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02325 2025-10-06 cs.CR cs.AI 57%

Agentic-AI Healthcare: Multilingual, Privacy-First Framework with MCP Agents

Mohammed A. Shehab

机构 * Concordia Continuing Education(康科德继续教育) Concordia University(康科德大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 1 figure. Submitted as a system/vision paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09650 2025-10-06 cs.CV cs.LG cs.MM cs.RO eess.IV 57%

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen, Junwei Zheng, Yufan Chen, Kailun Yang, Jiamin Wu, Chongqing Hao, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) HEBUST

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02534 2025-10-06 cs.SE 50%

ZeroFalse: Improving Precision in Static Analysis with LLMs

Mohsen Iranmanesh, Sina Moradi Sabet, Sina Marefat, Ali Javidi Ghasr, Allison Wilson, Iman Sharafaldin, Mohammad A. Tayebi

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他推理 4 篇

2510.03204 2025-10-06 cs.CL 57%

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar, Xing Han Lù, Léo Boisvert, Massimo Caccia, Jérémy Espinas, Alexandre Aussem, Véronique Eglin, Alexandre Lacoste

机构 * LIRIS - CNRS, INSA Lyon, Universite Claude Bernard Lyon 1(LIRIS - CNRS,INSA里昂,克劳德·贝尔纳大学里昂) Esker ServiceNow Research(ServiceNow研究) Mila - Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02752 2025-10-06 cs.CL 57%

The Path of Self-Evolving Large Language Models: Achieving Data-Efficient Learning via Intrinsic Feedback

Hangfan Zhang, Siyuan Xu, Zhimeng Guo, Huaisheng Zhu, Shicheng Liu, Xinrun Wang, Qiaosheng Zhang, Yang Chen, Peng Ye, Lei Bai, Shuyue Hu

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Singapore Management University(新加坡管理大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21117 2025-10-06 cs.IR cs.AI 57%

A Comprehensive Review on Harnessing Large Language Models to Overcome Recommender System Challenges

Rahul Raja, Anshaj Vats, Arpita Vats, Anirban Majumder

机构 * Linkedin, Carnegie Mellon University, Stanford University(LinkedIn、卡内基梅隆大学、斯坦福大学) Meta Linkedin, Meta AI, Amazon, Boston University(LinkedIn、Meta AI、亚马逊、波士顿大学) Amazon(亚马逊)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04404 2025-10-06 cs.AI 57%

LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers

Jingze Zhu, Yongliang Wu, Wenbo Zhu, Jiawang Cao, Yanqiang Zheng, Jiawei Chen, Xu Yang, Bernt Schiele, Jonas Fischer, Xinting Hu

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

Comments The submission was made before undergoing the required review by the co-authors' affiliated institutions. We are withdrawing the paper to allow for the completion of the institutional review process

详情

展开后加载摘要…

URL PDF HTML 收藏