arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-24 至 2025-10-24 共收录 13 信号源:cs.CL, cs.AI, cs.LG

1. 复杂问题求解 13 篇

2505.14667 2025-10-24 cs.AI cs.CL 87%

SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment

Wonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert No

机构 * Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学) Department of Computer Science and Engineering, Yonsei University(计算机科学与工程系,延世大学)

专题命中 复杂问题求解 :reasoning(title,abstract);chain-of-thought(title);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025. Code and models are available at https://ai-isl.github.io/safepath

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22651 2025-10-24 cs.CV cs.CL cs.LG 86%

Sherlock: Self-Correcting Reasoning in Vision-Language Models

Yi Ding, Ruqi Zhang

机构 * Department of Computer Science, Purdue University, USA(计算机科学系,普渡大学)

专题命中 复杂问题求解 :reasoning(title,abstract);CoT(abstract);self-correction(abstract);分类 cs.CL、cs.LG

Comments Published at NeurIPS 2025, 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18254 2025-10-24 cs.AI cs.LG 84%

Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning

Sion Weatherhead, Flora Salim, Aaron Belbasis

机构 * University of New South Wales Sydney(新南威尔士大学悉尼分校) Aurecon Group(Aurecon集团)

专题命中 复杂问题求解 :reasoning(title,abstract);self-correction(abstract);分类 cs.AI、cs.LG

Comments Currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20780 2025-10-24 cs.CL cs.AI 81%

Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost

Runzhe Zhan, Zhihong Huang, Xinyi Yang, Lidia S. Chao, Min Yang, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20607 2025-10-24 cs.LG cs.AI 81%

Generalizable Reasoning through Compositional Energy Minimization

Alexandru Oarga, Yilun Du

机构 * University of Barcelona(巴塞罗那大学) Harvard University(哈佛大学)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20188 2025-10-24 cs.AI 79%

TRUST: A Decentralized Framework for Auditing Large Language Model Reasoning

Morris Yu-Chao Huang, Zhen Tan, Mohan Zhang, Pingzhi Li, Zhuo Zhang, Tianlong Chen

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04462 2025-10-24 cs.CL cs.AI 73%

Benchmarking GPT-5 for biomedical natural language processing

Yu Hou, Zaifu Zhan, Min Zeng, Yifan Wu, Shuang Zhou, Rui Zhang

专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20653 2025-10-24 stat.ML cs.AI cs.LG 66%

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

Jack Butler, Nikita Kozodoi, Zainab Afolabi, Brian Tyacke, Gaiar Baimuratov

机构 * Amazon Web Services(亚马逊网络服务) Zalando

专题命中 复杂问题求解 :reasoning(abstract,journal_ref);分类 cs.AI、cs.LG

Journal ref Neural Information Processing Systems (NeurIPS 2025) Workshop: Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20807 2025-10-24 cs.CV cs.LG 57%

Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers

Dean L Slack, G Thomas Hudson, Thomas Winterbottom, Noura Al Moubayed

机构 * Durham University(杜伦大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.LG

Comments 14 pages, 14 figures

Journal ref IEEE Transactions on Neural Networks and Learning Systems, 36, 19106-19118, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20637 2025-10-24 cs.LG 57%

Large Multimodal Models-Empowered Task-Oriented Autonomous Communications: Design Methodology and Implementation Challenges

Hyun Jong Yang, Hyunsoo Kim, Hyeonho Noh, Seungnyun Kim, Byonghyo Shim

机构 * Seoul National University(首尔国立大学) Hanbat National University(翰baum国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11851 2025-10-24 cs.CR cs.CL 57%

Deep Research Brings Deeper Harm

Shuo Chen, Zonggen Li, Zhen Han, Bailan He, Tong Liu, Haokun Chen, Georg Groh, Philip Torr, Volker Tresp, Jindong Gu

机构 * LMU Munich(慕尼黑大学) Siemens(西门子) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Technical University of Munich (TUM)(慕尼黑技术大学) AWS AI(亚马逊AI) Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse卓越可靠AI学校) University of Hong Kong (HKU)(香港大学) University of Oxford(牛津大学)

专题命中 复杂问题求解 :planning(abstract);分类 cs.CL

Comments Accepted to Reliable ML from Unreliable Data Workshop @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20622 2025-10-24 cs.CV 50%

SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding

Yuan Sheng, Yanbin Hao, Chenxu Li, Shuo Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学)

专题命中 复杂问题求解 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20223 2025-10-24 cs.CR cs.MM 50%

Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations

Divyanshu Kumar, Shreyas Jena, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

专题命中 复杂问题求解 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏