arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-23 至 2025-12-23 共收录 5 信号源:cs.CL, cs.AI, cs.LG

1. 逻辑推理 5 篇

2512.19092 2025-12-23 cs.CL 89%

A Large Language Model Based Method for Complex Logical Reasoning over Knowledge Graphs

基于大型语言模型的复杂知识图谱逻辑推理方法

Ziyan Zhang, Chao Wang, Zhuo Chen, Lei Chen, Chiyi Li, Kai Song

机构 * School of Information Science and Engineering, Chongqing Jiaotong University(信息科学与工程学院,重庆交通大学) State Grid Chongqing Electric Power Company(国网重庆市电力公司)

专题命中 逻辑推理 :reasoning(title,abstract);logical reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 本文提出ROG方法,结合KG邻域检索与LLM推理,有效提升复杂知识图谱逻辑推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13704 2025-12-23 cs.CV 85%

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

TiViBench: 用于视频生成模型的视频推理基准测试

Harold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu, Zihan Wang, Sirui Chen, Wenkai Cheng, Kanghao Chen, Hongfei Zhang, Zixin Zhang, Rongjin Guo, Yu Cheng, Ying-Cong Chen

专题命中 逻辑推理 :reasoning(title,abstract);planning(abstract);logical reasoning(abstract)

AI总结 TiViBench提出用于评估视频生成模型的推理能力,结合VideoTPO提升推理性能,揭示商业与开源模型在推理潜力上的差异。

Comments Project: https://haroldchen19.github.io/TiViBench-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17920 2025-12-23 cs.CL cs.AI 62%

Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression

分离约束合规性与语义准确性:一种新的基准,用于在压缩下评估指令遵循

Rahul Baxi

机构 * Independent Researcher(独立研究者)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CDCT基准,揭示LLMs在压缩下约束合规性与语义准确性之间的矛盾,发现中等压缩时约束违规主要由RLHF训练的有用性行为导致。

Comments 19 pages, 9 figures; currently under peer review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19466 2025-12-23 cs.CY cs.CL cs.HC 57%

Epistemological Fault Lines Between Human and Artificial Intelligence

人类与人工智能之间的知识论断层

Walter Quattrociocchi, Valerio Capraro, Matjaž Perc

机构 * Department of Computer Science, Sapienza University of Rome, Rome, Italy Department of Psychology, University of Milan Bicocca, Milan, Italy Faculty of Natural Sciences Mathematics, University of Maribor, Maribor, Slovenia Community Healthcare Center Dr. Adolf Drolc Maribor, Maribor, Slovenia University College, Korea University, Seoul, Republic of Korea Department of Physics, Kyung Hee University, Seoul, Republic of Korea

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL

AI总结 本文揭示大型语言模型与人类认知在知识生成机制上的结构性差异,指出LLMs是随机模式完成系统而非知识代理,并识别七种知识断层,对社会评估、治理及知识素养提出影响。

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18177 2025-12-23 cs.AI cs.CV 57%

NEURO-GUARD: Neuro-Symbolic Generalization and Unbiased Adaptive Routing for Diagnostics -- Explainable Medical AI

NEURO-GUARD:神经符号泛化与无偏自适应路由用于诊断——可解释的医疗AI

Midhat Urooj, Ayan Banerjee, Sandeep Gupta

机构 * Arizona State University(亚利桑那州立大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI

AI总结 NEURO-GUARD通过整合视觉变换器与语言驱动推理,提升医疗图像诊断的准确性、透明性和泛化能力。

Comments Accepted at Asilomar Conference

详情

展开后加载摘要…

URL PDF HTML 收藏