arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-01 至 2025-12-01 共收录 26 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 26 篇

2505.14684 2025-12-01 cs.CL cs.AI 88%

Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning

注意间隔:为改进链式推理微调而弥合思维跳跃

Haolei Xu, Yuchen Yan, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Shengpei Jiang, Kaitao Song, Weiming Lu, Jun Xiao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) SF Technology(SF技术) Microsoft Research Asia(微软亚洲研究院)

专题命中 推理评测 :chain-of-thought(title,abstract);CoT(abstract,comments);reasoning(abstract);logical reasoning(abstract)

AI总结 本文提出CoT思维跳跃桥任务,通过生成缺失的中间推理步骤提升数学推理任务的性能,实验表明在NuminaMath上改进达+5.87%。

Comments Accepted to NeurIPS 2025. Camera ready version. Code: https://github.com/ZJU-REAL/Mind-the-Gap Project: https://zju-real.github.io/CoT-Bridge/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03408 2025-12-01 cs.CL 85%

Efficient Reasoning via Thought-Training and Thought-Free Inference

通过思维训练和无思维推理实现高效推理

Canhui Wu, Qiong Cao, Chao Xue, Wei Xi, Xiaodong He

机构 * Xi’an Jiaotong University(西安交通大学) JD Future Academy(京东未来学院)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

AI总结 3TF通过训练混合模型实现高效推理,使模型能在无显式推理生成的情况下进行高质量推理。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07984 2025-12-01 cs.CV 85%

SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing

SAMChat:引入链式推理和GRPO以增强小规模遥感遥感小语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中东部技术大学(METU)电子与电气工程系)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 SAMChat通过引入链式推理和GRPO,专为遥感影像分析优化,实现了在开放描述和分类任务上的高精度表现。

Comments Accepted to Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS) Special Issue on Foundation and Large Vision Models for Remote Sensing. Code and dataset are available at https://github.com/aybora/SAMChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01943 2025-12-01 cs.CL cs.AI cs.CV cs.RO 81%

ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks

ROVER:基于视觉语言模型的视频递归推理用于具身任务

Philip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo, James Glass

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) RAI Institute(RAI研究院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 ROVER通过递归分解视频轨迹提升具身任务中的视频推理能力,有效减少幻觉并提高准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01584 2025-12-01 cs.AI cs.LG 81%

ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models

ReasoningWeekly: 一个面向大语言模型的通用知识和逻辑推理挑战

Zixuan Wu, Francesca Lucchetti, Aleksander Boruch-Gruszecki, Jingmiao Zhao, Carolyn Jane Anderson, Joydeep Biswas, Federico Cassano, Arjun Guha

机构 * Northeastern University(东北大学) Wellesley College(韦尔斯利学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) Cursor

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 ReasoningWeekly是一个基于NPR周日谜题挑战的通用知识和逻辑推理基准,揭示了现有评估中不明显的模型能力差距,并发现了新的推理失败类型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23269 2025-12-01 cs.AI 79%

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

OctoMed:面向尖端多模态医疗推理的数据配方

Timothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin, Reuben Tan, Tristan Naumann, Junjie Hu, Hoifung Poon

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Microsoft Research(微软研究院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 OctoMed通过结构化推理轨迹的数据配方,提升医疗多模态推理模型的性能和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23112 2025-12-01 cs.CV cs.LG 79%

MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?

MathSight: 一个探索视觉语言模型在大学级数学推理中是否真正看到的基准

Yuandong Wang, Yao Cui, Yuxin Zhao, Zhen Yang, Yangfu Zhu, Zhenzhou Shao

机构 * Capital Normal University(首都师范大学) Tsinghua University(清华大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

AI总结 MathSight基准测试通过对比不同视觉输入条件,探索视觉语言模型在大学级数学推理中视觉信息的真实贡献。

Comments Comments: 32 pages, 15 figures, 9 tables, includes appendix. Project page: https://cnu-bot-group.github.io/MathSight/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23477 2025-12-01 cs.CV 78%

Video-CoM: Interactive Video Reasoning via Chain of Manipulations

Video-CoM:通过操作链进行交互式视频推理

Hanoona Rasheed, Mohammed Zumri, Muhammad Maaz, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) University of California Merced(加州梅尔 Ced 大学) Google Research(谷歌研究院) Linköping University(林奈大学) Australian National University(澳大利亚国立大学)

专题命中 推理评测 :reasoning(title,abstract)

AI总结 Video-CoM通过操作链实现交互式视频推理,提升视频理解任务的性能和可解释性

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22134 2025-12-01 cs.CV cs.RO 78%

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action

DualVLA: 通过推理与行动部分解耦构建通用具身代理

Zhen Fang, Zhuoyang Liu, Jiaming Liu, Hao Chen, Yu Zeng, Shiting Huang, Zehui Chen, Lin Chen, Shanghang Zhang, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院) CUHK(香港大学)

专题命中 推理评测 :reasoning(title,abstract)

AI总结 DualVLA通过推理与行动部分解耦,提升通用具身代理的行动与多模态理解平衡能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22055 2025-12-01 cs.CV cs.MM 75%

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

OralGPT-Omni: 一种多功能的牙科多模态大语言模型

Jing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan, Wenkai Zhou, Kaixin Guo, Zanting Ye, Yanpeng Sun, Xinyu Zhang, Yanqi Yang, Qiankun Li, Hao Tang, James Kit-Hon Tsoi, Linlin Shen, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) School of Biomedical Engineering, Southern Medical University(南方医科大学生物医学工程学院) Singapore University of Technology and Design(新加坡科技与设计大学) University of Auckland(奥克兰大学) University of Science and Technology of China(中国科学技术大学) School of Computer Science, Peking University(北京大学计算机学院) College of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 OralGPT-Omni是一种专门用于牙科的多模态大语言模型,通过TRACE-CoT数据集和四阶段训练范式,实现了对牙科图像的高效分析和高准确率评估。

Comments 47 pages, 42 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21757 2025-12-01 cs.CY cs.AI cs.CL cs.CR 62%

Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs

医疗恶意:用于医疗LLM中情境感知安全的数据库

Andrew Maranhão Ventura D'addario

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 Medical Malice提出一个情境感知安全的医疗数据库,通过对抗性提示和伦理推理提升医疗LLM的安全性,以应对医疗领域复杂且系统性的威胁。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10157 2025-12-01 cs.AI cs.CL 62%

One Patient, Many Contexts: Scaling Medical AI with Contextual Intelligence

一个患者,许多情境:通过情境智能扩展医疗AI

Michelle M. Li, Ben Y. Reis, Adam Rodman, Tianxi Cai, Noa Dagan, Ran D. Balicer, Joseph Loscalzo, Isaac S. Kohane, Marinka Zitnik

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出情境切换技术,通过调整模型推理而非重新训练,使医疗AI能适应不同专科、人群和地理环境,提升其在现实护理中的可靠性和扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23387 2025-12-01 cs.AI 57%

Hierarchical AI-Meteorologist: LLM-Agent System for Multi-Scale and Explainable Weather Forecast Reporting

分层人工智能气象学家:用于多尺度和可解释性天气预报报告的LLM代理系统

Daniil Sukhorukov, Andrei Zakharov, Nikita Glazkov, Katsiaryna Yanchanka, Vladimir Kirilin, Maxim Dubovitsky, Roman Sultimov, Yuri Maksimov, Ilya Makarov

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出分层人工智能气象学家,通过多尺度推理和关键词生成提升天气预报的可解释性和鲁棒性。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19079 2025-12-01 cs.CV cs.AI cs.HC 57%

Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition

读微笑:面部情绪识别基础模型中的代理偏差

Iosif Tsangko, Andreas Triantafyllopoulos, Adem Abdelmoula, Adria Mallol-Ragolta, Bjoern W. Schuller

机构 * CHI – Chair of Health Informatics, MRI, Technical University of Munich(慕尼黑技术大学健康信息学系) MCML – Munich Center for Machine Learning(慕尼黑机器学习中心) GLAM – Group on Language, Audio, & Music(语言、音频与音乐小组) MDSI – Munich Data Science Institute(慕尼黑数据科学研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文研究了面部情绪识别基础模型中视觉线索的依赖性及潜在偏差,揭示了模型在情绪推理中的内在一致性及公平性风险。

Journal ref IEEE Access, Early Access, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01734 2025-12-01 cs.CL 57%

Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs

本福特定律的诅咒:追踪数字偏差到大语言模型中的数值幻觉

Jiandong Shao, Yao Lu, Jianfei Yang

机构 * Nanyang Technological University(南洋理工大学) University College London(伦敦大学学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 研究发现LLMs在预训练中学习到的数字偏倚导致数值生成偏差,通过修剪特定神经元可缓解错误输出。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22275 2025-12-01 cs.AI 57%

RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender Systems

RecToM:用于评估基于LLM的对话推荐系统机器理论 of mind 的基准

Mengfan Li, Xuanhua Shi, Yang Deng

机构 * SMU(德克萨斯南方大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 RecToM是一个用于评估基于LLM的对话推荐系统中机器理论 of mind 能力的新基准,重点关注认知推理和行为预测两个维度。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21982 2025-12-01 cs.CV cs.AI 57%

DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models

DialBench: 通过大基础模型实现指针表盘读数的准确识别

Futian Wang, Chaoliu Weng, Xiao Wang, Zhen Chen, Zhicheng Zhao, Jin Tang

机构 * School of Computer Science and Technology, Anhui University, Hefei 230601, China(安徽大学计算机科学与技术学院) School of Artificial Intelligence, Anhui University, Hefei 230601, China(安徽大学人工智能学院) Department of Computer Science and Information Technology, La Trobe University, Bendigo, Australia(拉筹伯大学计算机科学与信息技术系)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出DialBench基准数据集和MRLM模型,通过物理关系注入提升指针表盘读数识别的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19957 2025-12-01 cs.CL 57%

AppSelectBench: Application-Level Tool Selection Benchmark

AppSelectBench: 应用级工具选择基准

Tianyi Chen, Michael Solodko, Sen Wang, Jongwoo Ko, Junheng Hao, Colby Banbury, Sara Abdali, Saeed Amizadeh, Qing Xiao, Yinheng Li, Tianyu Ding, Kamran Ghasedi Dizaji, Suzhen Zheng, Hao Fan, Justin Wagle, Pashmina Cameron, Kazuhito Koishida

机构 * Microsoft(微软)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 AppSelectBench通过评估CUAs的应用选择能力,揭示了模型在跨应用推理中的系统性优劣,为智能代理的研究提供了新基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19220 2025-12-01 cs.CV cs.AI 57%

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

大视觉语言模型真的在医学图像上具有基础性吗?来自意大利临床视觉问答的证据

Federico Felizzi, Olivia Riccomi, Michele Ferramola, Francesco Andrea Causio, Manuel Del Medico, Vittorio De Vita, Lorenzo De Mori, Alessandra Piscitelli, Pietro Eric Risuleo, Bianca Destro Castaniti, Antonio Cristiano, Alessia Longo, Luigi De Angelis, Mariapia Vassalli, Marcello Di Pumpo

机构 * SIIAM NSBProject Dept. of Life Sciences & Public Health, UCSC(生命科学与公共卫生系,UCSC) ASL RM 4 UCSC Univ. Paris Cité(巴黎Cité大学) Univ. of Pisa(比萨大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 研究通过测试四种先进模型在意大利医学问题上的表现,揭示了大视觉语言模型在视觉基础上的差异,强调了临床部署前的严格评估需求。

Comments Accepted at the Workshop on Multimodal Representation Learning for Healthcare (MMRL4H), EurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14307 2025-12-01 cs.CL 57%

How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs

大语言模型如何在叙述中理解时间意义:一项对大语言模型认知评估的案例研究

Karin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu, Khanh Chi Le, Michael Mensink, Ahn Thu Tong, Dongyeop Kang

机构 * University of Minnesota(明尼苏达大学) Hamline University(哈姆林大学) University of Wisconsin-Stout(威斯康星州立大学斯托特分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本研究通过实验发现,大语言模型在处理叙述中的时间意义时,过度依赖原型性,导致不一致的判断和因果推理困难,表明其与人类在理解叙述方面存在根本差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23220 2025-12-01 cs.CV 50%

Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day

针对表格数据生成的大型语言模型指令微调——一天内完成

Milad Abdollahzadeh, Abdul Raheem, Zilong Zhao, Uzair Javaid, Kevin Yee, Nalam Venkata Abhishek, Tram Truong-Huu, Biplab Sikdar

机构 * SIT(新加坡科技学院) NUS(新加坡国立大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 本文提出了一种在有限资源下通过指令微调提升LLM表格数据生成能力的方法,利用高质量数据集和少量指令实现与GPT-4o相当的生成性能。

Comments Accepted International Conference on Machine Learning (ICML 2025), 1st Workshop on Foundation Models for Structured Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06754 2025-12-01 cs.RO cs.CV 50%

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

SlotVLA:迈向机器人操作中物体-关系表示的建模

Taisei Hanyu, Nhat Chung, Huy Le, Toan Nguyen, Yuki Ikebe, Anthony Gunderman, Duy Nguyen Ho Minh, Khoa Vo, Tung Kieu, Kashu Yamazaki, Chase Rainwater, Anh Nguyen, Ngan Le

机构 * University of Arkansas(亚拉巴马大学) FPT Software AI Center(FPT软件人工智能中心) University of Stuttgart(斯图加特大学) Aalborg University(奥尔堡大学) Carnegie Mellon University(卡内基梅隆大学) University of Liverpool(利物浦大学) German Research Center for Artificial Intelligence(德国人工智能研究中心) Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究学校)

专题命中 推理评测 :reasoning(abstract)

AI总结 SlotVLA通过引入LIBERO+数据集和基于槽注意力的框架,实现高效且可解释的机器人操作中物体-关系表示建模。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06224 2025-12-01 cs.CV 50%

TEFormer: Texture-Aware and Edge-Guided Transformer for Semantic Segmentation of Urban Remote Sensing Images

TEFormer:面向城市遥感图像语义分割的纹理感知与边缘引导变换器

Guoyu Zhou, Jing Zhang, Yi Yan, Hui Zhang, Li Zhuo

专题命中 推理评测 :planning(abstract)

AI总结 TEFormer通过纹理感知和边缘引导的变换器提升城市遥感图像语义分割的精度与鲁棒性。

Comments Accepted by IEEE GRSL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22256 2025-12-01 cs.CV 50%

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

UMind-VL:一种通用的超声视觉-语言模型,用于统一的 grounded perception 和全面的 interpretation

Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao, Junjie Hou, Yaqian Wang, Nianxi Liao, Anlan Sun, Fei Gao, Jia Ding, Yuhang Liu, Dong Wang

机构 * Yizhun Medical AI Team(义诊医疗AI团队)

专题命中 推理评测 :reasoning(abstract)

AI总结 UMind-VL 是一种通用超声视觉-语言模型,通过统一的 grounded perception 和 comprehensive interpretation 实现对医学影像的高效理解和诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20330 2025-12-01 cs.RO cs.CV 50%

ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation

ArtiBench 和 ArtiBrain:可泛化的视觉-语言操纵基准测试

Yuhan Wu, Tiantian Wei, Shuo Wang, ZhiChao Wang, Yanyong Zhang, Daniel Cremers, Yan Xia

机构 * University of Science and Technology of China(中国科学技术大学) Technical University of Munich(慕尼黑技术大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 ArtiBench和ArtiBrain通过统一高层推理与自适应低层控制,提升可操纵物体操作的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18842 2025-12-01 cs.CV 50%

Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset

通过大规模多模态数据集增强描述性图像质量评估

Zhiyuan You, Jinjin Gu, Xin Cai, Zheyuan Li, Kaiwen Zhu, Chao Dong, Tianfan Xue

机构 * The Chinese University of Hong Kong(香港中文大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Sofia University(索菲亚大学) University of Macau(澳门大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 本研究提出DepictQA-Wild模型,通过构建大规模多模态数据集提升图像质量评估的准确性和实用性。

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏