arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-24 至 2025-10-24 共收录 81 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 17 篇

2510.20603 2025-10-24 cs.AI cs.CL 81%

What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation

Heejin Do, Jaehui Hwang, Dongyoon Han, Seong Joon Oh, Sangdoo Yun

机构 * ETH Zürich, ETH AI Center(苏黎世联邦理工学院,ETH人工智能中心) NAVER AI Lab(NAVER人工智能实验室) University of Tübingen(图宾根大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18458 2025-10-24 cs.CL cs.AI cs.CV 81%

Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning

Wenyi Xiao, Leilei Gan

机构 * Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19258 2025-10-24 cs.CL cs.AI 81%

Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning

Yu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong, Yue Dong, Wen Xiao

机构 * University of California, Riverside(加州大学河滨分校) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Microsoft(微软公司)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ICLR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19836 2025-10-24 cs.AI cs.SY eess.SY 79%

Benchmarking Reasoning Reliability in Artificial Intelligence Models for Energy-System Analysis

Eliseo Curcio

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20193 2025-10-24 cs.IR cs.CL cs.CV cs.LG 76%

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.LG

Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25271 2025-10-24 cs.AI cs.CV cs.LG cs.MA 62%

RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration

Xiuyuan Chen, Jian Zhao, Yuchen Yuan, Tianle Zhang, Huilin Zhou, Zheng Zhu, Ping Hu, Linghe Kong, Chi Zhang, Weiran Huang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) University of Science and Technology of China(中国科学技术大学) GigaAI School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13938 2025-10-24 cs.LG cs.AI cs.LO cs.PL cs.SE 62%

CLEVER: A Curated Benchmark for Formally Verified Code Generation

Amitayush Thakur, Jasper Lee, George Tsoukalas, Meghana Sistla, Matthew Zhao, Stefan Zetzsche, Greg Durrett, Yisong Yue, Swarat Chaudhuri

机构 * UT Austin(得克萨斯大学) Amazon(亚马逊) Caltech(加州理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01243 2025-10-24 cs.CV cs.AI cs.CL 62%

Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants

Lixiong Qin, Shilong Ou, Miaoxuan Zhang, Jiangning Wei, Yuhang Zhang, Xiaoshuai Song, Yuchen Liu, Mei Wang, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 推理评测 :CoT(abstract);分类 cs.CL、cs.AI

Comments 50 pages, 14 figures, 42 tables. NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19892 2025-10-24 cs.CL cs.AI 62%

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

Nishant Balepur, Dang Nguyen, Dayeon Ki

机构 * University of Maryland(马里兰大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted as a Spotlight paper at the EMNLP 2025 Wordplay Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20337 2025-10-24 cs.AI 57%

Collateral Damage Assessment Model for AI System Target Engagement in Military Operations

Clara Maathuis, Kasper Cools

机构 * Open University of the Netherlands(荷兰开放大学) Royal Military Academy, Belgium(比利时皇家军事学院) Vrije Universiteit Brussel, Belgium(比利时自由大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Accepted at MILCOM 2025 WS07

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20119 2025-10-24 cs.LG 57%

There is No "apple" in Timeseries: Rethinking TSFM through the Lens of Invariance

Arian Prabowo, Flora D. Salim

机构 * School of Computer Science and Engineering(计算机科学与工程学院) University of New South Wales(新南威尔士大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19167 2025-10-24 cs.CL 57%

"You Are Rejected!": An Empirical Study of Large Language Models Taking Hiring Evaluations

Dingjie Fu, Dianxing Shi

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Technical Report, 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15086 2025-10-24 cs.CL 57%

Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models

Yeonjun In, Wonjoong Kim, Kanghoon Yoon, Sungchul Kim, Mehrab Tanjim, Sangwu Park, Kibum Kim, Chanyoung Park

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20113 2025-10-24 eess.SY cs.SD cs.SY 50%

SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance

Haowei Lou, Chengkai Huang, Hye-young Paik, Yongquan Hu, Aaron Quigley, Wen Hu, Lina Yao

机构 * University of New South Wales(新南威尔士大学) Macquarie University(麦考瑞大学) National University of Singapore(国立新加坡大学) CSIRO's Data61(澳大利亚联邦科学与工业研究组织Data61部门)

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16793 2025-10-24 cs.CV 50%

REOBench: Benchmarking Robustness of Earth Observation Foundation Models

Xiang Li, Yong Tao, Siyuan Zhang, Siwei Liu, Zhitong Xiong, Chunbo Luo, Lu Liu, Mykola Pechenizkiy, Xiao Xiang Zhu, Tianjin Huang

机构 * University of Bristol, UK(英国布里斯托大学) University of Exeter, UK(英国埃克塞特大学) South China Normal University, China(华南师范大学) The University of Aberdeen, UK(英国阿伯丁大学) Technical University of Munich, Germany(慕尼黑技术大学) Eindhoven University of Technology, NL(埃因霍温理工大学)

专题命中 推理评测 :planning(abstract)

Comments Accepted to NeruIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 6 篇

2510.19851 2025-10-24 cs.CR cs.AI 89%

Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability

Artur Zolkowski, Wen Xing, David Lindner, Florian Tramèr, Erik Jenner

机构 * ML Alignment & Theory Scholars (MATS)(机器学习对齐与理论学者(MATS)) ETH Zurich(苏黎世联邦理工学院) Google DeepMind(谷歌DeepMind)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11570 2025-10-24 cs.CR cs.CL 79%

Bag of Tricks for Subverting Reasoning-based Safety Guardrails

Shuo Chen, Zhen Han, Haokun Chen, Bailan He, Shengyun Si, Jingpei Wu, Philip Torr, Volker Tresp, Jindong Gu

机构 * LMU Munich(慕尼黑大学) Siemens(西门子) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Technical University of Berlin(柏林技术大学) Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse可靠性人工智能卓越学院) DFKI AWS AI(AWS人工智能) University of Oxford(牛津大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL

Comments OpenAI Red-teaming Challenge Winner and Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15293 2025-10-24 cs.LG cs.AI 62%

LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models

Qianyue Hao, Yiwen Song, Qingmin Liao, Jian Yuan, Yong Li

机构 * Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系、北京理工大学、清华大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20369 2025-10-24 cs.LG 57%

Ask a Strong LLM Judge when Your Reward Model is Uncertain

Zhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu, Ilgee Hong, Changlong Yu, Wenlin Yao, Yao Liu, Haoming Jiang, Lihong Li, Hyokun Yun, Tuo Zhao

机构 * Georgia Institute of Technology(佐治亚理工学院) Amazon(亚马逊)

专题命中 其他推理 :reasoning(abstract);分类 cs.LG

Comments NeurIPS 2025, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17536 2025-10-24 cs.CL cs.MA 57%

Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?

Hyeong Kyu Choi, Xiaojin Zhu, Sharon Li

机构 * Department of Computer Sciences, University of Wisconsin-Madison(计算机科学系,威斯康星大学麦迪逊分校)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21401 2025-10-24 cs.CV 50%

JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation

Md Jueal Mia, M. Hadi Amini

机构 * Knight Foundation School of Computing and Information Sciences (KFSCIS)(骑士基金会计算与信息科学学院) Florida International University(佛罗里达国际大学) Sustainability, Optimization, and Learning for InterDependent networks laboratory (solid lab)(可持续性、优化与互依赖网络学习实验室)

专题命中 其他推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏