arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-08-18 至 2025-08-18 共收录 13 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 13 篇

2508.09666 2025-08-18 cs.CL 85%

Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation

Ziyang Ma, Qingyue Yuan, Linhai Zhang, Deyu Zhou

专题命中 推理评测 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11021 2025-08-18 cs.CV cs.CL 79%

Can Multi-modal (reasoning) LLMs detect document manipulation?

Zisheng Liang, Kidus Zewde, Rudra Pratap Singh, Disha Patil, Zexi Chen, Jiayu Xue, Yao Yao, Yifei Chen, Qinzhe Liu, Simiao Ren

机构 * Duke University(杜克大学) Indian Institute of Technology, Roorkee(印度理工学院,罗尔基) New York University(纽约大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Wisconsin Madison(威斯康星大学麦迪逊分校) Columnbia University(哥伦比亚大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments arXiv admin note: text overlap with arXiv:2503.20084

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18932 2025-08-18 cs.MM cs.CL 79%

MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks

Lei Zhang, Xin Zhou, Chaoyue He, Di Wang, Yi Wu, Hong Xu, Wei Liu, Chunyan Miao

机构 * Nanyang Technological University(南洋理工大学) University College London(伦敦大学学院) Alibaba Group(阿里巴巴集团)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11305 2025-08-18 cs.SE 78%

Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning

Xin Wang, Zhenhao Li, Zishuo Ding

专题命中 推理评测 :reasoning(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10947 2025-08-18 cs.CV 78%

MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text

Ronghao Xu, Zhen Huang, Yangbo Wei, Xiaoqian Zhou, Zikang Xu, Ting Liu, Zihang Jiang, S. Kevin Zhou

专题命中 推理评测 :reasoning(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11318 2025-08-18 cs.CL 57%

LLM Compression: How Far Can We Go in Balancing Size and Performance?

Sahil Sk, Debasish Dhal, Sonal Khosla, Sk Shahid, Sambit Shekhar, Akash Dhaka, Shantipriya Parida, Dilip K. Prasad, Ondřej Bojar

机构 * Odia Generative AI, India(奥迪生成人工智能,印度) AMD Silo AI, Finland(AMD Silo AI,芬兰) The Arctic University of Norway, Norway(挪威北极大学,挪威) Charles University, MFF, ÚFAL, Czech Republic(查尔斯大学,捷克共和国)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments This paper has been accepted for presentation at the RANLP 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11163 2025-08-18 cs.CL 57%

MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering

Hikaru Asano, Hiroki Ouchi, Akira Kasuga, Ryo Yonetani

机构 * The University of Tokyo \ AIP Japan Tokyo Nara Institute of Science The University of Tokyo \ AIP

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10976 2025-08-18 cs.AI 57%

Grounding Rule-Based Argumentation Using Datalog

Martin Diller, Sarah Alice Gaggl, Philipp Hanisch, Giuseppina Monterosso, Fritz Rauschenbach

机构 * Logic Programming and Argumentation Group, TU Dresden, Germany(图灵编程与论证组,德累斯顿理工大学,德国) Knowledge-Based Systems Group, TU Dresden, Germany(知识系统组,德累斯顿理工大学,德国) DIMES - University of Calabria, Italy(迪梅斯-卡拉布里亚大学,意大利)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10972 2025-08-18 cs.CV cs.AI cs.HC 57%

Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision

Rosiana Natalie, Wenqian Xu, Ruei-Che Chang, Rada Mihalcea, Anhong Guo

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09057 2025-08-18 cs.CL 57%

MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions

Zeyu Huang, Juyuan Wang, Longfeng Chen, Boyi Xiao, Leng Cai, Yawen Zeng, Jin Xu

机构 * South China University of Technology(华南理工大学) Pazhou Lab(Pazhou 实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06813 2025-08-18 cs.LG cs.PL 57%

Technical Report: Full-Stack Fine-Tuning for the Q Programming Language

Brendan R. Hogan, Will Brown, Adel Boyarsky, Anderson Schneider, Yuriy Nevmyvaka

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11153 2025-08-18 cs.CV 50%

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

Maoquan Zhang, Bisser Raytchev, Xiujuan Sun

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(Hiroshima大学研究生院) Department of Computer Science, Weifang University of Science and Technology(潍坊科技大学计算机科学系)

专题命中 推理评测 :reasoning(abstract)

Comments The International Conference on Neural Information Processing (ICONIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 50%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 推理评测 :reasoning(abstract)

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏