arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5878 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5878 篇

2511.06345 2025-11-25 cs.DC cs.AI 57%

PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization

PRAGMA:一种基于剖析的多智能体框架用于自动内核优化

Kelun Lei, Hailong Yang, Huaitao Zhang, Xin You, Kaige Zhang, Zhongzhi Luan, Yi Liu, Depei Qian

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 PRAGMA通过整合执行反馈和硬件剖析,利用AI生成高性能内核,实现比现有方法更优的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09755 2025-11-25 cs.CE cs.AI 57%

Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era

智能设计4.0:迈向代理AI时代的范式演变

Shuo Jiang, Min Xie, Frank Youhua Chen, Jian Ma, Jianxi Luo

机构 * Department of Systems Engineering(系统工程系) City University of Hong Kong(香港城市大学) Department of Decision Analytics and Operations(决策分析与运营系) Department of Information Systems(信息系统系)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出智能设计4.0范式,基于基础模型的代理AI系统,推动工程设计的自动化与智能化发展。

Comments 17 pages, 7 figures

Journal ref ASME Journal of Computing and Information Science in Engineering, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18397 2025-11-25 cs.AI cs.SE 57%

Natural Emergent Misalignment from Reward Hacking in Production RL

生产强化学习中奖励黑客导致的自然涌现偏差

Monte MacDiarmid, Benjamin Wright, Jonathan Uesato, Joe Benton, Jon Kutasov, Sara Price, Naia Bouscal, Sam Bowman, Trenton Bricken, Alex Cloud, Carson Denison, Johannes Gasteiger, Ryan Greenblatt, Jan Leike, Jack Lindsey, Vlad Mikulik, Ethan Perez, Alex Rodrigues, Drake Thomas, Albert Webson, Daniel Ziegler, Evan Hubinger

机构 * Anthropic

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 研究发现大型语言模型在生产强化学习环境中学习奖励黑客会导致严重偏差,提出三种缓解措施以减少这种偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17515 2025-11-25 cs.HC cs.AI 57%

Embedding Generative AI into Systems Analysis and Design Curriculum: Framework, Case Study, and Cross-Campus Empirical Evidence

将生成式人工智能融入系统分析与设计课程:框架、案例研究和跨校园实证证据

Mahmoud Elkhodr, Ergun Gide

机构 * School of Engineering and Technology, Central Queensland University(工程与技术学院,中央昆士兰大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出SAGE框架,通过将生成式人工智能融入课程设计,培养学生批判性思维和AI编排能力,揭示学生在系统分析中的关键挑战和教育改进方向。

Comments ~12,000 words; 4 figures; 6 tables; multi-site study (across 4 Australian campuses)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14169 2025-11-25 cs.CV cs.AI 57%

AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs

AdaTok: 一种基于对象感知表示的自适应令牌压缩方法,用于高效多模态大语言模型

Xinliang Zhang, Lei Zhu, Hangzhou He, Shuang Zeng, Ourui Fu, Jiakui Hu, Zhengjian Yao, Yanye Lu

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学医学部医学技术研究所) Department of Biomedical Engineering, Peking University(北京大学生物医学工程系) National Biomedical Imaging Center, Peking University(北京大学国家生物医学成像中心)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 AdaTok通过自适应令牌压缩方法,利用对象感知表示提升多模态大语言模型的效率,实现高压缩比与高性能的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07334 2025-11-25 cs.CL 57%

Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation

框架表示假说:多词LLM可解释性与概念引导的文本生成

Pedro H. V. Valois, Lincon S. Souza, Erica K. Shimomoto, Kazuhiro Fukui

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 本文提出框架表示假说,通过多词建模提升LLM的可解释性,利用概念引导解码实现更安全透明的文本生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16691 2025-11-24 cs.CL 57%

Reproducibility Report: Test-Time Training on Nearest Neighbors for Large Language Models

可重复性报告:为大语言模型进行邻近邻居测试时训练

Boyang Zhou, Johan Lindqvist, Lindsey Li

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 通过测试时训练邻近邻居方法,提升大语言模型在多种领域上的性能,尤其在结构化数据集上表现显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16402 2025-11-21 cs.AI cs.DB 57%

Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance

可信AI在代理湖仓中的实现:从并发到治理

Jacopo Tagliabue, Federico Bianchi, Ciro Greco

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出Bauplan设计,通过事务机制实现湖仓中的数据和计算隔离,解决代理工作流的可信性问题,并提供自修复管道的实现。

Comments AAAI26, pre-print of paper accepted at the Trustworthy Agentic AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16229 2025-11-21 cs.CR cs.AI 57%

Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security

Q-MLLM:向量量化用于鲁棒多模态大语言模型安全

Wei Zhao, Zhe Li, Yige Li, Jun Sun

机构 * Singapore Management University(新加坡管理大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 Q-MLLM通过两级向量量化提升多模态大语言模型的安全性,有效防御对抗性攻击并保持模型实用性。

Comments Accepted by NDSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16004 2025-11-21 cs.SE cs.AI 57%

InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

InfCode: 基于对抗的测试和补丁迭代精炼框架用于可靠的软件问题解决

KeFan Li, Mengfei Wang, Hengzhi Zhang, Zhichao Li, Yuan Yuan, Mu Li, Xiang Gao, Hailong Sun, Chunming Hu, Weifeng Lv

机构 * Beihang University(北航大学) Beijing Tokfinity Technology Co., Ltd.(北京 Tokfinity 技术有限公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 InfCode通过对抗多代理框架实现测试和补丁的迭代精炼,以提高软件问题解决的可靠性,实验表明其在SWE-bench Verified上达到79.4%的性能,创下了新的最先进水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15895 2025-11-21 cs.AI 57%

Decomposing Theory of Mind: How Emotional Processing Mediates ToM Abilities in LLMs

解构理论思维:情绪处理如何调节LLMs的理论思维能力

Ivan Chulo, Ananya Joshi

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究通过对比激活加法引导,发现LLMs的理论思维能力提升由情绪处理而非分析推理所驱动。

Comments Published at ToM4AI workshop@AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15712 2025-11-21 cs.CR cs.AI cs.DC 57%

Secure Autonomous Agent Payments: Verifying Authenticity and Intent in a Trustless Environment

安全的自主代理支付:在无信任环境中验证真实性和意图

Vivek Acharya

机构 * Vivek Acharya(独立研究者)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出基于区块链的框架,通过密码学认证和零知识证明,实现自主代理支付中的意图验证和身份认证,确保交易安全性和可追溯性。

Comments 6 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15378 2025-11-20 cs.AI 57%

Terra Nova: A Comprehensive Challenge Environment for Intelligent Agents

Trevor McInroe

机构 * The University of Edinburgh(爱丁堡大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14307 2025-11-19 cs.SD cs.LG 57%

Audio Question Answering with GRPO-Based Fine-Tuning and Calibrated Segment-Level Predictions

Marcel Gibier, Nolwenn Celton, Raphaël Duroselle, Pierre Serrano, Olivier Boeffard, Jean-François Bonastre

机构 * Inria Paris, LR2(巴黎Inria)

专题命中 其他推理 :reasoning(abstract);分类 cs.LG

Comments Submission to Track 5 of the DCASE 2025 Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18970 2025-11-19 cs.AI 57%

LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions

Xixun Lin, Yucheng Ning, Jingwen Zhang, Yan Dong, Yilong Liu, Yongxuan Wu, Xiaohua Qi, Nan Sun, Yanmin Shang, Kun Wang, Pengfei Cao, Qingyue Wang, Lixin Zou, Xu Chen, Chuan Zhou, Jia Wu, Peng Zhang, Qingsong Wen, Shirui Pan, Bin Wang, Yanan Cao, Kai Chen, Songlin Hu, Li Guo

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Nanyang Technological University(南洋理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hong Kong University of Science and Technology(香港科技大学) School of Cyber Science and Engineering, Wuhan University(武汉大学网络安全科学与工程学院) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学中关村人工智能学院) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) School of Computing, Faculty of Science and Engineering, Macquarie University(麦考瑞大学计算机学院) Cyberspace Institute of Advanced Technology, Guangzhou University(广州大学高级技术网络研究所) Squirrel Ai Learning Xiaomi Company(小米公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12782 2025-11-18 cs.CL cs.CR 57%

LLM Reinforcement in Context

Thomas Rivasseau

机构 * McGill University(麦吉尔大学)

专题命中 其他推理 :chain-of-thought(abstract);分类 cs.CL

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12579 2025-11-18 cs.AI 57%

Enhancing Conversational Recommender Systems with Tree-Structured Knowledge and Pretrained Language Models

Yongwen Ren, Chao Wang, Peng Du, Chuan Qin, Dazhong Shen, Hui Xiong

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00829 2025-11-18 cs.CL 57%

Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation

Yanming Sun, Runzhe Zhan, Chi Seng Cheang, Han Wu, Xuebo Liu, Yuyao Niu, Fengying Ye, Kaixin Lan, Lidia S. Chao, Derek F. Wong

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12638 2025-11-18 cs.CV cs.AI 57%

edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer

Chen Qian, Xinran Yu, Zewen Huang, Danyang Li, Qiang Ma, Fan Dang, Xuan Ding, Guangyong Shang, Zheng Yang

机构 * Tsinghua University(清华大学) Beijing Jiaotong University(北京交通大学) Inspur Yunzhou Industrial Internet Co., Ltd(Inspur云洲工业互联网有限公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12131 2025-11-18 cs.CV cs.AI 57%

OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description

Quanxing Xu, Ling Zhou, Feifei Zhang, Jinyu Tian, Rubing Huang

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11867 2025-11-18 cs.CL 57%

Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM Approaches

Namu Park, Giridhar Kaushik Ramachandran, Kevin Lybarger, Fei Xia, Ozlem Uzuner, Meliha Yetisgen, Martin Gunn

机构 * Department of Information Sciences and Technology, George Mason University, Fairfax, VA, USA(信息科学与技术系,乔治·玛莎大学,弗吉尼亚州, Fairfax) Department of Linguistics, University of Washington, Seattle, WA, USA(语言学系,华盛顿大学,西雅图,华盛顿州,美国) Department of Radiology, School of Medicine, University of Washington, Seattle, WA, USA(放射学系,医学院,华盛顿大学,西雅图,华盛顿州,美国)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments Submitted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11126 2025-11-17 cs.CL cs.CV 57%

Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion

Yi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei, Wenzhi Chen

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14009 2025-11-17 cs.CL 57%

Activation-Guided Consensus Merging for Large Language Models

Yuxuan Yao, Shuqi Liu, Zehua Liu, Qintong Li, Mingyang Liu, Xiongwei Han, Zhijiang Guo, Han Wu, Linqi Song

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) City University of Hong Kong Shenzhen Research Institute(香港城市大学深圳研究院) Huawei Noah’s Ark Lab, Hong Kong SAR(华为诺亚实验室(香港特别行政区)) University of Hong Kong(香港大学) Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10552 2025-11-14 cs.CL 57%

URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding

Yongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng, Zening Lin, Lianwen Jin

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments Accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19594 2025-11-13 cs.CL 57%

Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs

Jun Bai, Minghao Tong, Yang Liu, Zixia Jia, Zilong Zheng

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26550 2025-11-12 cs.AI 57%

EdgeRunner 20B: Military Task Parity with GPT-5 while Running on the Edge

Jack FitzGerald, Aristotelis Lazaridis, Dylan Bates, Aman Sharma, Jonnathan Castillo, Yousif Azami, Sean Bailey, Jeremy Cao, Peter Damianov, Kevin de Haan, Luke Kerbs, Vincent Lu, Joseph Madigan, Jeremy McLaurin, Jonathan Tainer, Dave Anderson, Jonathan Beck, Jamie Cuticello, Colton Malkerson, Tyler Saltsman

机构 * Anonymous Authors(匿名作者)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

Comments 19 pages; v2 includes an additional appendix with test set examples

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07480 2025-11-12 cs.CR cs.AI 57%

KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs

Shuyuan Liu, Jiawei Chen, Xiao Yang, Hang Su, Zhaoxia Yin

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07427 2025-11-12 cs.DC cs.AI 57%

DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones

Tuowei Wang, Minxing Huang, Fengzu Li, Ligeng Chen, Jinrui Zhang, Ju Ren

机构 * Tsinghua University(清华大学) Honor Device Co., Ltd.(荣耀设备有限公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07097 2025-11-11 cs.AI cs.MA 57%

Agentic AI Sustainability Assessment for Supply Chain Document Insights

Diego Gosmar, Anna Chiara Pallotta, Giovanni Zenezini

机构 * Head of AI, Tesisquare(Tesisquare人工智能负责人) Voiceinteroperability.ai Initiative Member(Voiceinteroperability.ai项目成员) Linux Foundation AI & Data(Linux基金会人工智能与数据) Functional Analyst, Tesisquare(Tesisquare功能分析师) MSc in Engineering Management Polytechnic University of Turin(理工学院工程管理硕士) Assistant Professor in PM and Supply Chain Management Polytechnic University of Turin(理工学院项目管理与供应链管理副教授)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06991 2025-11-11 cs.LG 57%

CoLM: Collaborative Large Models via A Client-Server Paradigm

Siqi Huang, Sida Huang, Hongyuan Zhang

专题命中 其他推理 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏