arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-30 至 2025-12-30 共收录 87 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 24 篇

2512.22245 2025-12-30 cs.LG cs.AI 62%

Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation

校准大语言模型法官:用于快速可靠不确定性估计的线性探针

Bhaktipriya Radharapu, Eshika Saxena, Kenneth Li, Chenxi Whitehouse, Adina Williams, Nicola Cancedda

机构 * FAIR at Meta(Meta 的 FAIR 研究所) Meta Superintelligence Labs(Meta 超智能实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种基于线性探针的校准方法,通过Brier分数损失训练,实现快速可靠的LLM法官不确定性估计,相比现有方法在计算效率和泛化能力上表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23029 2025-12-30 cs.DC cs.AI 57%

Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware

私有LLM服务器的可行性和性能:基于Qwen3-30B在消费级硬件上的基准分析

Alex Khalil, Guillaume Heilles, Maria Parraga, Simon Heilles

机构 * UCLouvain(列日大学) Universidad Espíritu Santo(圣灵大学) DENEM Labs(DENEM实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文基于Qwen3-30B在消费级硬件上评估私有LLM服务器的可行性和性能,证明其在成本和隐私方面对SMBs的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22899 2025-12-30 cs.AI cs.CV 57%

HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery

HiSciBench: 一个分层多学科基准,用于从阅读到发现的科学智能

Yaping Zhang, Qixuan Zhang, Xingquan Zhang, Zhiyuan Chen, Wenwen Zhuang, Yupu Liang, Lu Xiang, Yang Zhao, Jiajun Zhang, Yu Zhou, Chengqing Zong

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 HiSciBench是一个分层多学科基准,用于评估从阅读到发现的科学智能,涵盖五个层次,揭示了模型在不同科学推理阶段的能力差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12411 2025-12-30 cs.CL 57%

The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases

大语言模型的文化基因:跨语料库训练对模型价值观与偏见的影响研究

Emanuel Z. Fenech-Borg, Tilen P. Meznaric-Kos, Milica D. Lekovic-Bojovic, Arni J. Hentze-Djurhuus

机构 * Department of Communications(通讯系) University of Malta(马耳他大学) Faculty of Mathematics(数学系) University of Primorska(普里摩尔卡大学) Faculty of Electrical Engineering(电气工程系) University of Montenegro(黑山大学) Faculty of Science & Technology(科学与技术系) University of the Faroe Islands(法罗群岛大学) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本研究通过跨文化探针数据集分析大语言模型在个人主义-集体主义和权力距离维度上的文化偏见,揭示模型价值观与训练语料文化背景的关联。

Comments 10 pages, 5 figures, IEEE conference format, submitted to [Conference Name]

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10516 2025-12-30 cs.CV cs.AI 57%

CogStream: Context-guided Streaming Video Question Answering

CogStream: 基于上下文的流视频问答

Zicheng Zhao, Kangyu Wang, Shijie Li, Rui Qian, Weiyao Lin, Huabin Liu

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出CogStream任务,通过视觉流压缩和历史对话检索,解决流视频中关键信息识别与问答问题。

Comments Project page: https://github.com/LiamZhao326/CogStream

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19466 2025-12-30 cs.CV cs.LG 57%

ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection

ForgerySleuth: 赋能多模态大语言模型进行图像篡改检测

Zhihao Sun, Haoran Jiang, Haoran Chen, Yixin Cao, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 ForgerySleuth通过多模态大语言模型进行图像篡改检测,利用线索融合和数据集构建提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22737 2025-12-30 cs.CL 57%

WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference

WeDLM:弥合扩散语言模型与标准因果注意力以实现快速推理

Aiwei Liu, Minghua He, Shaoxun Zeng, Sijun Zhang, Linhao Zhang, Chuhan Wu, Wei Jia, Yuan Liu, Xiao Zhou, Jie Zhou

机构 * WeChat AI, Tencent(腾讯微信AI) Peking University(北京大学) Tsinghua University(清华大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 WeDLM通过使用标准因果注意力实现高效并行解码,显著提升推理速度,优于优化的AR引擎。

Comments 23 pages, 8 figures, project page: https://wedlm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22629 2025-12-30 cs.AI cs.IR 57%

DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation

DICE:基于概率评分的离散可解释比较评估用于检索增强生成

Shiyan Liu, Jian Ma, Rui Qu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 DICE提出一种基于概率评分的两阶段框架,提升RAG评估的可解释性和稳健性,通过透明判断和系统性错误诊断,实现高效且可信的评估。

Comments Accepted at ResponsibleFM @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22496 2025-12-30 cs.MA cs.AI 57%

Hierarchical Pedagogical Oversight: A Multi-Agent Adversarial Framework for Reliable AI Tutoring

层级教学监督:一种多智能体对抗框架用于可靠的AI辅导

Saisab Sadhu, Ashim Dhor

机构 * AAAI 2026 EGSAI Community Activity(AAAI 2026 EGSAI 社区活动)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出层级教学监督框架,通过多智能体对抗机制提升AI辅导的可靠性与效率,实验结果显示其在资源受限环境下表现优异。

Comments Accepted for presentation at the AAAI 2026 EGSAI Community Activity (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20663 2025-12-30 cs.MA cs.AI cs.SY eess.SY 57%

MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems

MTTR-A:多智能体系统中认知恢复延迟的测量

Barak Or

机构 * Office of the CEO, MetaOr Artificial Intelligence(MetaOr人工智能首席执行官办公室) Google–Reichman Tech School, Reichman University(Reichman大学Google–Reichman技术学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出MTTR-A指标,用于测量多智能体系统中认知恢复延迟,通过理论分析和实验验证建立了运行时认知可靠性的量化基础。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13204 2025-12-30 cs.CV 50%

RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection

RefineVAD: 语义引导的特征重校准用于弱监督视频异常检测

Junhee Lee, ChaeBeen Bang, MyoungChul Kim, MyeongAh Cho

专题命中 推理评测 :reasoning(abstract)

AI总结 RefineVAD通过结合时间动态和语义结构,利用双过程推理提升弱监督视频异常检测的性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05526 2025-12-30 cs.CV 50%

When Deepfake Detection Meets Graph Neural Network:a Unified and Lightweight Learning Framework

当深度伪造检测遇见图神经网络:一种统一且轻量级的学习框架

Haoyu Liu, Chaoyu Gong, Mengke He, Jiate Li, Kai Han, Siqiang Luo

机构 * Nanyang Technological University(南洋理工大学) University of Southern California(南加州大学) The University of Hong Kong(香港大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 本文提出SSTGNN,一种统一且轻量级的深度伪造检测框架,通过图神经网络联合处理空间、时间及频谱信息,实现高效且准确的伪造检测。

Comments Accepted to KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22218 2025-12-30 cs.CV cs.MM 50%

Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark

面向路牌的视觉问答:ViSignVQA数据集、方法与基准

Hieu Minh Nguyen, Tam Le-Thanh Dang, Kiet Van Nguyen

机构 * University of Information Technology, Ho Chi Minh City, Vietnam(胡志明市信息技术大学) Vietnam National University(越南国家大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 ViSignVQA首次提出大规模越南语路牌多模态数据集,结合OCR与预训练语言模型,提升低资源语言VQA性能。

Comments Dataset paper; code and data will be released

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22197 2025-12-30 cs.CV 50%

Quadrant Segmentation VLM with Few-Shot Adaptation and OCT Learning-based Explainability Methods for Diabetic Retinopathy

四象限分割视觉语言模型结合少样本适应与基于OCT学习的可解释方法用于糖尿病视网膜病变

Shivum Telang

机构 * Department of Biostatics(生物统计学系) North Allegheny Senior High School(北阿勒格尼高中) University of Pittsburgh Rangos Research Center(匹兹堡大学Rangos研究中心)

专题命中 推理评测 :reasoning(abstract)

AI总结 本文提出一种基于四象限分割和少样本学习的多模态视觉语言模型,利用OCT和视网膜图像生成可解释性热图,提升糖尿病视网膜病变的诊断效果。

Comments 4 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20933 2025-12-30 cs.SE 50%

Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code

大型语言模型代码设计能力的分层评估

Mootez Saad, Boqi Chen, José Antonio Hernández López, Dániel Varró, Tushar Sharma

专题命中 推理评测 :reasoning(abstract)

AI总结 本文评估了大型语言模型在代码设计能力上的分层表现,发现其在耦合性推理上易受噪声影响,而内聚性分析在受控环境下较稳健,但整体自主推理能力有限。

Comments 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 12 篇

2512.21257 2025-12-30 cs.IR cs.CL 83%

ReaSeq: Unleashing World Knowledge via Reasoning for Sequential Modeling

ReaSeq:通过推理解锁世界知识用于序列建模

Jiakai Tang, Chuan Wang, Gaoming Yang, Han Wu, Jiahao Yu, Jian Wu, Jianwu Hu, Junjun Zheng, Longbin Li, Shuwen Xiao, Xiangheng Kong, Yeqiu Yang, Yuning Jiang, Ahjol Nurlanbek, Binbin Cao, Bo Zheng, Fangmei Zhu, Gaoming Zhou, Huimin Yi, Huiping Chu, Jin Huang, Jinzhe Shan, Kenan Cui, Longbin Li, Silu Zhou, Wen Chen, Xia Ming, Xiang Gao, Xin Yao, Xingyu Wen, Yan Zhang, Yiwen Hu, Yulin Wang, Ziheng Bao, Zongyuan Wu

机构 * TaoRank Team(TaoRank团队)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 ReaSeq通过引入世界知识增强推理,提升推荐系统在物品表示和用户兴趣建模上的性能,实现IPV、CTR、订单和GMV的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13797 2025-12-30 cs.CL 79%

Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons

痕迹推理:通过压缩信标实现的内存高效推理

Giovanni Monea, Yair Feldman, Shankar Padmanabhan, Kianté Brantley, Yoav Artzi

机构 * Cornell University(康奈尔大学) Harvard University(哈佛大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL

AI总结 本研究提出通过压缩信标实现内存高效推理,利用联合蒸馏和强化学习框架优化缓存压缩,提升大语言模型在长上下文推理中的内存与准确性平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07583 2025-12-30 cs.CL cs.AI 62%

Complementary Learning Approach for Text Classification using Large Language Models

基于大语言模型的文本分类互补学习方法

Navid Asgari, Benjamin M. Cole

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种基于大语言模型的文本分类互补学习方法,通过人机协作弥补各自弱点,以低成本技术处理评分差异问题。

Comments After further review, we identified substantive issues that materially affect the validity of the manuscript's core results and conclusions. Addressing these would require a fundamental reworking of the analysis and framing. To maintain the integrity of the public record, we request withdrawal of this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01956 2025-12-30 cs.AI cs.LG cs.MA 62%

Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models

通过多智能体语言模型实现临床笔记中临床级特征生成的扩展

Jiayi Wang, Jacqueline Jil Vallon, Nikhil V. Kotha, Neil Panjwani, Xi Ling, Margaret Redfield, Sushmita Vij, Sandy Srinivas, John Leppert, Mark K. Buyyounouski, Mohsen Bayati

机构 * Department of Management Science and Engineering, Stanford University School of Engineering(管理科学与工程系,斯坦福大学工程学院) Department of Radiation Oncology, Stanford University School of Medicine(放射肿瘤学系,斯坦福大学医学院) Operations, Information and Technology, Stanford University Graduate Business School(运营、信息与技术,斯坦福大学商学院) Graduate Business School Research Hub, Stanford University Graduate Business School(商学院研究中心,斯坦福大学商学院) Department of Medicine (Oncology), Stanford University School of Medicine(医学系(肿瘤学),斯坦福大学医学院) Department of Medicine, Stanford University School of Medicine(医学系,斯坦福大学医学院) Department of Urology, Stanford University School of Medicine(泌尿学系,斯坦福大学医学院) Veterans Affairs Palo Alto Health Care System(退伍军人事务帕洛阿尔托医疗系统) Department of Electrical Engineering, Stanford University School of Engineering(电气工程系,斯坦福大学工程学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出了一种多智能体语言模型系统,通过自动化临床笔记特征生成,实现了与人工方法相当的预测性能,并在不同医疗场景中展示了良好的可扩展性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15759 2025-12-30 cs.CL cs.AI cs.CV 62%

Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs

视觉增强大语言模型:赋能大语言模型中的多模态知识存储与共享

Yunxin Li, Zhenyu Liu, Baotian Hu, Wei Wang, Yuxin Ding, Xiaochun Cao, Min Zhang

机构 * Research Institute of Computing and Intelligence(计算与智能研究 institute) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MKS2方法,通过多模态知识存储与共享增强大语言模型的推理能力,提升其在物理和常识知识场景下的表现。

Comments 21 pages, 7 figures; Accepted by IEEE TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22508 2025-12-30 cs.LG cs.AI 62%

Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals

利用元数据和幻觉信号预测牙科修复学中大语言模型的正确性

Lucky Susanto, Anasta Pranawijayana, Cortino Sukotjo, Soni Prasad, Derry Wijaya

机构 * 1 Department of Data Science, Monash University Indonesia, Tangerang, Indonesia 2 Independent Researcher 3 Department of Prosthodontics, University of Pittsburgh, Pittsburgh, Pennsylvania 4 Department of Restorative Sciences, University of North Carolina Adams School of Dentistry, Chapel Hill, North Carolina 5 Department of Computer Science, Boston University, Boston, Massachusetts

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文研究通过元数据和幻觉信号预测牙科修复学中LLM的正确性,发现元数据方法可提升准确性,但需进一步改进以适应高风险应用。

Comments Accepted as a Short Paper at HEALTHINF2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23480 2025-12-30 cs.CR cs.AI 57%

Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation

面向软件供应链安全的代理AI:超越溯源到漏洞缓解

Toqeer Ali Syed, Mohammad Riyaz Belgaum, Salman Jan, Asadullah Abdullah Khan, Saad Said Alqahtani

机构 * Faculty of Computer and Information System(计算机与信息系统学院) Islamic University of Madinah(麦地那伊斯兰大学) Faculty of Computer Studies(计算机研究学院) Arab Open University-Bahrain(巴林阿拉伯开放大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出基于代理AI的软件供应链安全框架,结合LLM推理、强化学习和多代理协调,实现主动漏洞缓解,提升检测准确率和响应效率。

Comments Conference paper, accept in ACCA IEEE Bahrain

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23430 2025-12-30 cs.CL 57%

C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs

C2PO:诊断和解构大语言模型中的偏见捷径

Xuan Feng, Bo An, Tianlong Gu, Liang Chang, Fengrui Hao, Peipeng Yu, Shuai Zhao

机构 * Jinan University(济南大学) Nanyang Technological University(南洋理工大学) Engineering Research Center of Trustworthy AI (Ministry of Education)(可信人工智能工程研究中心) Guangxi Key Laboratory of Trusted Software(广西可信软件重点实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 C2PO通过因果对比偏好优化框架,解决大语言模型中的偏见问题,同时保持推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18186 2025-12-30 cs.SD cs.CL eess.AS 57%

Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought

通过上下文感知和推理链引导语言模型实现稳定的语音情感识别

Zhixian Zhao, Xinfa Zhu, Xinsheng Wang, Shuiyuan Wang, Xuelong Geng, Wenjie Tian, Lei Xie

专题命中 其他推理 :CoT(abstract);分类 cs.CL

AI总结 C$^2$SER通过上下文感知和推理链提升语音情感识别的稳定性和准确性,优于现有模型。

Comments This work has been published in IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01386 2025-12-30 cs.CL cs.CR cs.IR 57%

Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models

Topic-FlipRAG: 面向主题的对抗性观点操控攻击用于检索增强生成模型

Yuyang Gong, Zhuo Chen, Jiawei Liu, Miaokun Chen, Fengchang Yu, Wei Lu, Xiaofeng Wang, Xiaozhong Liu

机构 * Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) Worcester Polytechnic Institute(沃思堡理工学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

AI总结 本文提出Topic-FlipRAG,一种针对检索增强生成模型的面向主题对抗性观点操控攻击方法,通过两阶段流程影响模型输出观点,揭示了RAG系统安全防护的迫切需求。

Comments Accepted by USENIX Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23483 2025-12-30 cs.CV 50%

TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding

TV-RAG:一种具有时间意识和语义熵权的长视频检索与理解框架

Zongsheng Cao, Yangfan He, Anran Liu, Feng Chen, Zepeng Wang, Jun Xie

机构 * Researcher(研究者)

专题命中 其他推理 :reasoning(abstract)

AI总结 TV-RAG通过时间衰减检索和熵加权关键帧采样,提升长视频检索与理解性能,无需重新训练即可集成至现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01681 2025-12-30 physics.flu-dyn 50%

Large Language Model Driven Development of Turbulence Models

基于大语言模型的湍流模型开发

Zhongxin Yang, Yuanwei Bin, Yipeng Shi, Xiang I. A. Yang

专题命中 其他推理 :reasoning(abstract)

AI总结 本文提出利用大语言模型开发湍流模型,通过闭环迭代流程生成可解释且性能更优的近壁湍流模型,解决了不利压力梯度、系统旋转和表面粗糙度等问题。

Journal ref Flow (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏