arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-10 至 2025-12-10 共收录 51 信号源:cs.CL, cs.AI, cs.LG

1. 复杂问题求解 5 篇

2508.07871 2025-12-10 cs.CV 50%

CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning

CATP:面向高效增强多模态上下文学习的上下文自适应令牌剪枝

Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang

专题命中 复杂问题求解 :reasoning(abstract)

AI总结 CATP通过上下文自适应令牌剪枝方法,提升多模态上下文学习的效率和性能,减少冗余令牌带来的影响。

Comments 14 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 推理评测 11 篇

2512.08228 2025-12-10 cs.CV cs.AI 92%

MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models

MM-CoT:一种用于探测多模态模型中视觉链式推理的基准测试

Jusheng Zhang, Kaitong Cai, Xiaoyang Guo, Sidi Liu, Qinhan Lv, Ruiqi Chen, Jing Yang, Yijia Fan, Xiaofei Sun, Jian Wang, Ziliang Chen, Liang Lin, Keze Wang

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团) Snap Inc(Snap公司)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(title,abstract);logical reasoning(abstract)

AI总结 MM-CoT是一种用于评估多模态模型视觉链式推理能力的基准测试,通过验证推理链的视觉一致性和逻辑一致性,揭示生成模型在真实推理上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16205 2025-12-10 cs.AI 79%

ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025

ChemLabs on ChemO:一个用于IChO 2025多模态推理的多智能体系统

Qiang Xu, Shengyuan Bai, Leqing Chen, Zijing Liu, Yu Li

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 ChemLabs通过多智能体系统和SVE技术在ChemO基准测试中实现93.6分,推动化学问题自动解决的新进展。

Comments 13 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12200 2025-12-10 cs.AI cs.AR 79%

PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification

PRO-V-R1:基于推理增强的RTL验证编程代理

Yujie Zhao, Zhijing Wu, Boqin Yuan, Zhongming Yu, Hejia Zhang, Wentao Ni, Chia-Tung Ho, Haoxing Ren, Jishen Zhao

机构 * University of California San Diego(加州大学圣地亚哥分校) NVIDIA(NVIDIA公司)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 PRO-V-R1是一种基于推理增强的开源框架,通过结合LLM推理与编程工具,提升RTL验证的功能正确性和故障检测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20781 2025-12-10 cs.SE cs.AI 70%

Using LLMs in Generating Design Rationale for Software Architecture Decisions

在软件架构决策中使用LLMs生成设计理由

Xiyu Zhou, Ruiyin Li, Peng Liang, Beiqi Zhang, Mojtaba Shahin, Zengyang Li, Chen Yang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) RMIT University(皇家墨尔本理工大学) School of Computer Science, Central China Normal University(中央师范大学计算机学院) School of Artificial Intelligence, Shenzhen Polytechnic University(深圳职业技术学院人工智能学院)

专题命中 推理评测 :reasoning(abstract);CoT(abstract);分类 cs.AI

AI总结 本研究评估了LLMs在生成软件架构决策设计理由方面的性能,通过实验和访谈探讨了不同提示策略的效果及实际应用的可行性。

Comments Preprint accepted for publication in ACM Transactions on Software Engineering and Methodology (TOSEM), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14295 2025-12-10 cs.CL cs.AI cs.LG 67%

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

AraLingBench:一个用于评估大型语言模型阿拉伯语语言能力的人工标注基准

Mohammad Zbeeb, Hasan Abed Al Kader Hammoud, Sina Mukalled, Nadine Rizk, Fatima Karnib, Issam Lakkis, Ammar Mohanna, Bernard Ghanem

机构 * King Abdullah University of Science and Technology (KAUST)(卡斯泰克大学) American University of Beirut (AUB)(贝鲁特美国大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 AraLingBench通过150道人工标注的多项选择题评估阿拉伯语LLM的语言能力,揭示模型在语法和句法推理上的不足,强调了记忆与模式识别对模型性能的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07785 2025-12-10 physics.data-an cs.AI cs.LG hep-ex 62%

Automating High Energy Physics Data Analysis with LLM-Powered Agents

利用LLM代理自动化高能物理数据分析

Eli Gendreau-Distler, Joshua Ho, Dongwon Kim, Luc Tomas Le Pottier, Haichen Wang, Chengxi Yang

机构 * Department of Physics, University of California, Berkeley, Berkeley, CA 94720, USA(加州大学伯克利分校物理系) Physics Division, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA(伯克利国家实验室物理部)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本研究利用LLM代理自动化高能物理数据分析,通过混合系统结合LLM和Snakemake工作流管理器,评估代理在多阶段工作流中的性能。

Comments 16 pages, 6 figures, 2 tables, the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) - Machine Learning and the Physical Sciences (ML4PS) workshop (poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08936 2025-12-10 cs.AI cs.CL 62%

SimSUM: Simulated Benchmark with Structured and Unstructured Medical Records

SimSUM: 结构化和非结构化医疗记录的模拟基准

Paloma Rabaey, Stefan Heytens, Thomas Demeester

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 SimSUM通过模拟呼吸系统疾病患者记录,提供结构化与非结构化医疗数据的基准,用于研究临床信息提取及多模态数据生成。

Comments An earlier version of this dataset was published under the name SynSUM. It has since been renamed to SimSUM to avoid confusion with synthetic data generated from real data, and to emphasize the simulated nature of the dataset. The dataset is available at https://github.com/prabaey/SimSUM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08814 2025-12-10 cs.CL 57%

Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts

提问、回答与检测:基于角色扮演的LLM用于带有问题条件的专家混合模型的人格检测

Yifan Lyu, Liang Zhang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of International Business and Economics(国际经济贸易大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 ROME通过角色扮演LLM模拟用户回答心理测量问卷,生成可解释的证据链接语言线索与人格标签,提升人格检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07898 2025-12-10 cs.MA cs.AI 57%

MARINE: Theoretical Optimization and Design for Multi-Agent Recursive IN-context Enhancement

MARINE:多智能体递归上下文增强的理论优化与设计

Hongwei Zhang, Ji Lu, Yongsheng Du, Yanqin Gao, Lingjun Huang, Baoli Wang, Fang Tan, Peng Zou

机构 * Hongwei Zhang(张宏伟) Ji Lu(卢纪) Yongsheng Du(杜永生) Yanqin Gao(高燕琴) Lingjun Huang(黄令军) Baoli Wang(王宝利) Fang Tan(谭芳) Peng Zou(邹鹏)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 MARINE通过理论优化和设计实现多智能体递归上下文增强,显著提升推理性能并减少参数需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08233 2025-12-10 cs.RO 50%

Semantic-Metric Bayesian Risk Fields: Learning Robot Safety from Human Videos with a VLM Prior

语义-度量贝叶斯风险场:从人类视频中学习机器人安全性的方法

Timothy Chen, Marcus Dominguez-Kuhne, Aiden Swann, Xu Liu, Mac Schwager

机构 * Stanford University(斯坦福大学) California Institute of Technology(加州理工学院)

专题命中 推理评测 :planning(abstract)

AI总结 本文提出基于贝叶斯框架的语义-度量风险场,通过人类视频学习机器人安全风险模型,实现类人风险评估与规划。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07007 2025-12-10 cs.CV 50%

MELLM: A Flow-Guided Large Language Model for Micro-Expression Understanding

MELLM:一种面向微表情理解的流引导大型语言模型

Sirui Zhao, Zhengye Zhang, Shifeng Liu, Xinglong Mao, Shukang Yin, Chaoyou Fu, Tong Xu, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) Nanjing University(南京大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 MELLM通过结合光学流敏感性与LLM推理能力,首次实现对微表情的全面理解,显著提升微表情识别的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他推理 9 篇

2512.00218 2025-12-10 cs.AI cs.CR 88%

Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?

压力下的推理:训练激励如何影响推理链的可监控性?

Matt MacDermott, Qiyao Wei, Rada Djoneva, Francis Rhys Ward

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(title);CoT(abstract);分类 cs.AI

AI总结 本文研究了训练激励对推理链可监控性的影响,发现对抗性优化降低监控性能,而直接优化可监控性未显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08270 2025-12-10 cs.AI cs.CL q-fin.GN 81%

Reasoning Models Ace the CFA Exams

推理模型在CFA考试中表现优异

Jaisal Patel, Yunzhe Chen, Kaiwen He, Keyi Wang, David Li, Kairong Xiao, Xiao-Yang Liu

机构 * Rensselaer Polytechnic Institute(罗格斯理工学院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) SecureFinAI Lab(安全金融人工智能实验室) Columbia University(哥伦比亚大学) Business School(商学院)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文评估了先进推理模型在模拟CFA考试中的表现,发现Gemini 3.0 Pro在多个考试级别均取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08820 2025-12-10 cs.CV cs.AI 79%

Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning

无需训练的双双曲适配器用于更高效的跨模态推理

Yi Zhang, Chun-Wun Cheng, Junyi He, Ke Yu, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Aviles-Rivero

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Department of Electrical and Electronic Engineering, Southern University of Science and Technology(南方科技大学电子与电气工程系) Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学应用数学与理论物理系) Yau Mathematical Sciences Center, Tsinghua University(清华大学应用数学中心)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出无需训练的双双曲适配器方法,通过双曲空间嵌入提升跨模态推理性能,实现更高效的领域泛化和少样本识别。

Comments Accepted in IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10384 2025-12-10 cs.SI cs.AI cs.CL cs.CY 62%

Simulating Misinformation Propagation in Social Networks using Large Language Models

利用大语言模型模拟社交媒体上的虚假信息传播

Raj Gaurav Maurya, Vaibhav Shukla, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

机构 * Vizuara AI Labs(Vizuara AI实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文利用大语言模型模拟社交媒体虚假信息传播,通过构建人格代理网络研究虚假信息演变机制,揭示身份和意识形态驱动的人格加速虚假信息扩散,专家驱动人格则保持事实稳定。

Comments Accepted to CIKM 2025 Workshop LASS

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08524 2025-12-10 cs.CV cs.CL 57%

Beyond Real Weights: Hypercomplex Representations for Stable Quantization

超越真实权重:用于稳定量化 的超复数表示

Jawad Ibn Ahad, Maisha Rahman, Amrijit Biswas, Muhammad Rafsan Kabir, Robin Krambroeckers, Sifat Momen, Nabeel Mohammed, Shafin Rahman

机构 * Artificial Intelligence Department, RobotBulls Labs(机器人bulls实验室人工智能部门) Machine Intelligence Lab (MILab), North South University(北南大学机器智能实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 本文提出了一种基于超复数乘法的渐进式重新参数化策略,用于压缩多模态语言模型,实现参数和计算量的显著减少,同时保持模型性能。

Comments Accepted in Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08123 2025-12-10 cs.CL 57%

Universal Adversarial Suffixes Using Calibrated Gumbel-Softmax Relaxation

利用校准的Gumbel-Softmax松弛的通用对抗后缀

Sampriti Soor, Suklav Ghosh, Arijit Sur

机构 * Center for Intelligent Cyber Physical Systems(智能网络物理系统中心) Indian Institute of Technology Guwahati(印度古瓦哈提理工学院) Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 本文提出了一种利用校准的Gumbel-Softmax松弛学习通用对抗后缀的方法,能有效降低多种任务和模型的准确性。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07917 2025-12-10 cs.SE cs.AI physics.flu-dyn 57%

CFD-copilot: leveraging domain-adapted large language model and model context protocol to enhance simulation automation

CFD-copilot: 利用领域适应的大语言模型和模型上下文协议增强仿真自动化

Zhehao Dong, Shanghai Du, Zhen Lu, Yue Yang

机构 * State Key Laboratory for Turbulence and Complex Systems(湍流与复杂系统国家重点实验室) School of Mechanics and Engineering Science(力学与工程科学学院) Peking University(北京大学) HEDPS-CAPT

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 CFD-copilot通过领域适应的大语言模型和模型上下文协议提升CFD仿真的自动化水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14318 2025-12-10 physics.ed-ph cs.CY 50%

Report on the Scoping Workshop on AI in Science Education Research 2025

2025年人工智能在科学教育研究中的范围研讨会报告

Marcus Kubsch, Marit Kastaun, Peter Wulff, Nicole Graulich, Moriah Ariely, Alexander Bergmann-Gering, Sebastian Gombert, Bor Gregorcic, Hendrik Härtig, Benedikt Heuckmann, Andrea Horbach, Christina Krist, Gerd Kortemeyer, Ben Münch, Samuel Pazicni, Joshua M. Rosenberg, Sascha Schanze, Gena Sbeglia, Vidar Skogvoll, Christophe Speroni, Christoph Thyssen, Lars-Jochen Thoms, Brandon J. Yik, Xiaoming Zhai

专题命中 其他推理 :reasoning(abstract)

AI总结 2025年人工智能在科学教育研究中的范围研讨会报告,探讨AI在教育研究中的应用、挑战及负责任的整合方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06335 2025-12-10 cs.CV 50%

Harnessing Object Grounding for Time-Sensitive Video Understanding

利用物体接地提升时间敏感视频理解

Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi

机构 * Intel(英特尔公司)

专题命中 其他推理 :reasoning(abstract)

AI总结 本文提出 GO-Tokenizer 以提升视频大型语言模型的时间敏感视频理解能力,通过实时编码紧凑的物体信息,提高模型性能并减少噪声影响。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏