arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-22 至 2025-10-22 共收录 77 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 21 篇

2510.16263 2025-10-22 cs.RO cs.AI cs.CV 57%

NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?

Jierui Peng, Yanyan Zhang, Yicheng Duan, Tuo Liang, Vipin Chaudhary, Yu Yin

机构 * Department of Computer & Data Sciences(计算机与数据科学系) Case Western Reserve University(凯斯西储大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Homepage: https://vulab-ai.github.io/NEBULA-Alpha/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01042 2025-10-22 cs.LG cs.IR 57%

MatPROV: A Provenance Graph Dataset of Material Synthesis Extracted from Scientific Literature

Hirofumi Tsuruta, Masaya Kumagai

机构 * SAKURA internet Inc.(SAKURA互联网公司) Kyoto University(京都大学)

专题命中 推理评测 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11442 2025-10-22 cs.SE cs.LG 57%

ReVeal: Self-Evolving Code Agents via Reliable Self-Verification

Yiyang Jin, Kunzhao Xu, Hang Li, Xueting Han, Yanmin Zhou, Cheng Li, Jing Bai

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18873 2025-10-22 cs.CV 50%

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence

Ziang Zhang, Zehan Wang, Guanghao Zhang, Weilong Dai, Yan Xia, Ziang Yan, Minjie Hong, Zhou Zhao

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Shanghai AI Lab(上海人工智能实验室)

专题命中 推理评测 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18596 2025-10-22 cs.SE cs.CV 50%

CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent

Haojia Lin, Xiaoyu Tan, Yulei Qin, Zihan Xu, Yuchen Shi, Zongyi Li, Gang Li, Shaofei Cai, Siqi Cai, Chaoyou Fu, Ke Li, Xing Sun

机构 * Youtu-Agent Team(YouTu-Agent团队)

专题命中 推理评测 :reasoning(abstract)

Comments 24 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18262 2025-10-22 cs.CV 50%

UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding

Da Zhang, Chenggang Rong, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI)、中国电信)

专题命中 推理评测 :reasoning(abstract)

Comments We have released V1, which only reports the test results. Our work is still ongoing, and the next version will be coming soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21089 2025-10-22 cs.CV 50%

DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response

Junjue Wang, Weihao Xuan, Heli Qi, Zhihao Liu, Kunyi Liu, Yuhan Wu, Hongruixuan Chen, Jian Song, Junshi Xia, Zhuo Zheng, Naoto Yokoya

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Waseda University(早稻田大学) Stony Brook University(石溪大学) Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(abstract)

Comments A multi-hazard, multi-sensor, and multi-task vision-language dataset for global-scale disaster assessment and response

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 10 篇

2509.24248 2025-10-22 cs.AI cs.CL cs.LG 82%

SpecExit: Accelerating Large Reasoning Model via Speculative Exit

Rubing Yang, Huajun Bai, Song Liu, Guanghua Yu, Runzhi Fan, Yanbin Dang, Jiejing Zhang, Kai Liu, Jianchen Zhu, Peng Chen

机构 * Tencent(腾讯)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05188 2025-10-22 cs.CL cs.AI cs.LG math.ST stat.TH 82%

Counterfactual reasoning: an analysis of in-context emergence

Moritz Miller, Bernhard Schölkopf, Siyuan Guo

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ETH Zurich(苏黎世联邦理工学院) University of Cambridge(剑桥大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published as a conference paper at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06436 2025-10-22 cs.AI 79%

Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning

Song Yu, Xiaofei Xu, Ke Deng, Li Li, Lin Tian

机构 * School of Computer and Information Science, Southwest University(西南大学计算机与信息科学学院) School of Information Technology, Murdoch University(默多克大学信息科技学院) School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算技术学院) University of Technology Sydney(悉尼技术大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18425 2025-10-22 cs.AI 70%

Automated urban waterlogging assessment and early warning through a mixture of foundation models

Chenxu Zhang, Fuxiang Huang, Lei Zhang

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI

Comments Submitted to Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18502 2025-10-22 cs.CV cs.AI cs.CL cs.LG 67%

Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation

Wei-Chia Chang, Yan-Ann Chen

机构 * Yuan Ze University(元智大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by The 38th Conference of Open Innovations Association FRUCT, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01480 2025-10-22 cs.CV 67%

Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning

Kaihang Pan, Yang Wu, Wendong Bu, Kai Shen, Juncheng Li, Yingting Wang, Yunfei Li, Siliang Tang, Jun Xiao, Fei Wu, Hang Zhao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团)

专题命中 其他推理 :reasoning(abstract);CoT(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18297 2025-10-22 cs.CL cs.AI 62%

From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering

Lei Li, Xiao Zhou, Yingying Zhang, Xian Wu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Tencent Jarvis Lab(腾讯 Jarvis 实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17999 2025-10-22 cs.CY cs.AI cs.HC cs.LG 62%

The Narcissus Hypothesis: Descending to the Rung of Illusion

Riccardo Cadei, Christian Internò

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17937 2025-10-22 cs.LG cs.AI 62%

UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts

Fu-Yun Wang, Han Zhang, Michael Gharbi, Hongsheng Li, Taesung Park

机构 * Cuhk(香港中文大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17880 2025-10-22 cs.CL cs.AI 62%

Outraged AI: Large language models prioritise emotion over cost in fairness enforcement

Hao Liu, Yiqing Dai, Haotian Tan, Yu Lei, Yujia Zhou, Zhen Wu

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏