arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-01-22 至 2026-01-22 共收录 18 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 18 篇

2508.04339 2026-01-22 cs.AI 85%

Deliberative Reasoning Network: An Uncertainty-Driven Paradigm for Belief-Tracked Inference with Pretrained Language Models

辩证推理网络:一种基于不确定性的信念跟踪推理范式,用于预训练语言模型

Anran Xu, Jincheng Wang, Baigen Cai, Tao Wen

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(abstract);verifier(abstract);分类 cs.AI

AI总结 DRN通过不确定性最小化范式提升预训练语言模型的逻辑推理能力,实现高准确率和强泛化性能。

Comments This submission represents an early exploratory draft and was uploaded prematurely. The authors have decided to withdraw it because the current version does not accurately reflect the intended scope and technical formulation of the work, and may be misleading if cited

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15279 2026-01-22 cs.LG cs.AI 81%

MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs

MolecularIQ: 通过分子图上的符号验证来表征化学推理能力

Christoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl, Günter Klambauer, Sohvi Luukkonen

机构 * ELLIS Unit Linz and LIT AI Lab, Institute for Machine Learning, Johannes Kepler University, Linz, Austria(林茨ELLIS单元和LIT人工智能实验室,机器学习研究所,约翰·凯撒大学,林茨,奥地利) Clinical Research Institute Medical Artificial Intelligence, Johannes Kepler University, Linz, Austria(医学人工智能临床研究研究所,约翰·凯撒大学,林茨,奥地利)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 MolecularIQ是一个专注于分子图符号验证的化学推理基准测试,旨在评估模型对分子结构推理能力的细粒度评估并揭示模型在特定任务和分子结构上的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13761 2026-01-22 cs.AI cs.CL 81%

DARC: Decoupled Asymmetric Reasoning Curriculum for LLM Evolution

DARC:解耦的非对称推理课程用于LLM进化

Shengda Fan, Xuyan Ye, Yankai Lin

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 DARC通过解耦的非对称推理课程框架,提升大型语言模型的自我进化能力,实现无需人工标注的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08134 2026-01-22 cs.CL 79%

How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains

大型推理模型的置信度估计有多可靠?对高风险领域的系统基准测试

Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文通过系统基准测试,评估了大型推理模型置信度估计方法的可靠性,发现基于文本的编码器在歧视方面表现最佳,而结构感知模型在校准方面表现最佳,揭示了当前方法的局限性。

Comments Accepted to the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026) main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14700 2026-01-22 cs.CL 79%

DARL: Encouraging Diverse Answers for General Reasoning without Verifiers

DARL: 促进无验证者的一般推理中的多样化答案

Chongxuan Huang, Lei Lin, Xiaodong Shi, Wenping Hu, Ruiming Tang

机构 * School of Informatics, Xiamen University(厦门大学信息学院) Kuaishou Technology(快手科技) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),文化和旅游部)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 DARL通过鼓励在参考答案可控偏差范围内生成多样化答案,提升大语言模型在一般推理任务中的性能和输出多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14479 2026-01-22 cs.CL 79%

Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks

大语言模型的推理能力是否可信?:通过统计任务的人类基准测试

Crish Nagarkar, Leonid Bogachev, Serge Sharoff

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文通过统计任务的人类基准测试,评估了大语言模型在统计推理和自我评估能力上的表现,发现微调后的模型在统计任务上表现优异,且能更有效地评估答案质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14349 2026-01-22 cs.MA cs.LG 79%

MARBLE: Multi-Agent Reasoning for Bioinformatics Learning and Evolution

MARBLE: 多智能体推理用于生物信息学学习与进化

Sunghyun Kim, Seokwoo Yun, Youngseo Yun, Youngrak Lee, Sangsoo Lim

机构 * Division of AI Convergence, Dongguk University, Seoul, South Korea(东国大学人工智能融合系) AI Research Team, Ar-ge Inc., Seoul, South Korea(Ar-ge公司人工智能研究团队) Department of Biomedical Engineering, Dongguk University, Goyang-si, Gyeonggi-do, South Korea(东国大学生物医学工程系) Department of Computer Science and Artificial Intelligence, Dongguk University, Seoul, South Korea(东国大学计算机科学与人工智能系)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

AI总结 MARBLE通过多智能体推理实现生物信息学模型的稳定优化,提升性能并保持鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15115 2026-01-22 cs.CV 78%

Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning

无需训练的多阶段对抗推理 hateful 视频检测

Shuonan Yang, Yuchen Zhang, Zeyu Fu

机构 * Multimodal Intelligence Lab, Department of Computer Science, University of Exeter, United Kingdom(埃克塞特大学计算机科学系多模态智能实验室) Institute for Analytics and Data Science, University of Essex, United Kingdom(埃塞克斯大学分析与数据科学研究所)

专题命中 推理评测 :reasoning(title,abstract)

AI总结 MARS通过多阶段对抗推理框架实现无需训练的可解释仇恨视频检测,提升检测可靠性与透明度。

Comments Accepted at ICASSP 2026. \c{opyright} 2026 IEEE. This is the author accepted manuscript. The final published version will be available via IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14610 2026-01-22 cs.CV 78%

Learning Consistent Taxonomic Classification through Hierarchical Reasoning

通过层次推理学习一致的分类

Zhenghong Li, Kecheng Zheng, Haibin Ling

机构 * Stony Brook University(石溪大学) Ant Research(蚂蚁研究院)

专题命中 推理评测 :reasoning(title,abstract)

AI总结 VL-Taxon通过两阶段层次推理框架提升分类的叶级别准确性和层次一致性,实验表明其在iNaturalist-2021数据集上性能优于现有模型。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03369 2026-01-22 cs.CV cs.CL 74%

RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models

RiskCueBench: 视频语言模型中早期风险信号的前瞻性推理基准测试

Sha Luo, Yogesh Prabhu, Timothy Ossowski, Kaiping Chen, Junjie Hu

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of California San Diego(加州大学圣地亚哥分校)

专题命中 推理评测 :reasoning(title);分类 cs.CL

AI总结 RiskCueBench通过标注早期风险信号片段,评估视频语言模型在预测未来风险事件中的能力,揭示了现有系统在解读动态情境方面的不足。

Comments *updated author email in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15267 2026-01-22 cs.CY cs.AI cs.CL 62%

Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions

评估大型语言模型在法律应用中的表现:挑战、方法与未来方向

Yiran Hu, Huanghai Liu, Chong Wang, Kunran Li, Tien-Hsuan Wu, Haitao Li, Xinran Xu, Siqing Huo, Weihang Su, Ning Zheng, Siyuan Zheng, Qingyao Ai, Yun Liu, Renjun Bian, Yiqun Liu, Charles L. A. Clarke, Weixing Shen, Ben Kao

机构 * Tsinghua University(清华大学) The University of Hong Kong(香港大学) University of Waterloo(多伦多大学) Shanghai Jiaotong University(上海交通大学) Peking University(北京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨了大型语言模型在法律应用中的评估挑战与方法,分析了现有评估框架的局限性,并提出了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14695 2026-01-22 cs.LG cs.AI 62%

CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation

CoScale-RL: 通过协同扩展数据与计算实现高效后期训练

Yutong Chen, Jiandong Gao, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University, Beijing, China(清华大学电子工程系) College of AI, Tsinghua University, Beijing, China(清华大学人工智能学院) Beijing National Research Center for Information Science(北京信息科学国家研究中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 CoScale-RL通过协同扩展数据与计算,提升大型推理模型的训练效率和推理能力。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15282 2026-01-22 cs.CV cs.AI cs.RO 57%

Rethinking Video Generation Model for the Embodied World

重新思考具身世界的视频生成模型

Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, Daquan Zhou

机构 * Peking University(北京大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出Rbench机器人视频生成基准和RoVid-X数据集,旨在解决高保真机器人视频生成中的数据短缺问题,推动具身AI发展。

Comments Github: https://github.com/DAGroup-PKU/ReVidgen/ Project website: https://dagroup-pku.github.io/ReVidgen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15129 2026-01-22 cs.CL 57%

RSNA Large Language Model Benchmark Dataset for Chest Radiographs of Cardiothoracic Disease: Radiologist Evaluation and Validation Enhanced by AI Labels (REVEAL-CXR)

用于心胸疾病胸部X光影像的RSNA大语言模型基准数据集:通过AI标签增强的放射科评估和验证(REVEAL-CXR)

Yishu Wei, Adam E. Flanders, Errol Colak, John Mongan, Luciano M Prevedello, Po-Hao Chen, Henrique Min Ho Lee, Gilberto Szarf, Hamilton Shoji, Jason Sho, Katherine Andriole, Tessa Cook, Lisa C. Adams, Linda C. Chu, Maggie Chung, Geraldine Brusca-Augello, Djeven P. Deva, Navneet Singh, Felipe Sanchez Tijmes, Jeffrey B. Alpert, Elsie T. Nguyen, Drew A. Torigian, Kate Hanneman, Lauren K Groner, Alexander Phan, Ali Islam, Matias F. Callejas, Gustavo Borges da Silva Teles, Faisal Jamal, Maryam Vazirabad, Ali Tejani, Hari Trivedi, Paulo Kuriki, Rajesh Bhayana, Elana T. Benishay, Yi Lin, Yifan Peng, George Shih

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本研究通过AI辅助标注流程,创建了包含200张胸部影像的基准数据集,用于评估不同模型,同时帮助放射科医生更高效地标注影像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14903 2026-01-22 cs.CL 57%

PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation

PodBench: 一个全面的指令感知音频导向播客脚本生成基准

Chenning Xu, Mao Zheng, Mingyu Zheng, Mingyang Song

机构 * Large Language Model Department, Tencent(腾讯大语言模型部门)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 PodBench通过多方面评估框架和实验验证,揭示了开源模型在长上下文和多说话人协调任务中的鲁棒性优势,同时指出高指令遵循性与高质量内容之间的矛盾。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14549 2026-01-22 cs.LG 57%

QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design

QMC:通过异常感知量化和涌现记忆协同设计实现高效的SLM边缘推断

Nilesh Prasad Pandey, Jangseon Park, Onat Gungor, Flavio Ponzina, Tajana Rosing

机构 * University of California San Diego(加州大学圣地亚哥分校) San Diego State University(圣地亚哥州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 QMC通过异常感知量化和涌现记忆协同设计,实现高效SLM边缘推断,显著降低内存、能耗和延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14277 2026-01-22 cs.LG 57%

Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization on Llama-3.1-8B-Instruct

我应该使用哪种量化?对llama.cpp量化在Llama-3.1-8B-Instruct上的统一评估

Uygar Kurt

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 本文评估了llama.cpp不同量化方案在Llama-3.1-8B-Instruct模型上的性能,为选择合适的量化方法提供指导。

Comments 17 pages, 6 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11932 2026-01-22 cs.CV 50%

Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs

Hyperphantasia: 一个多模态大语言模型心理可视化能力评估基准

Mohammad Shahab Sepehri, Berk Tinaz, Zalan Fabian, Mahdi Soltanolkotabi

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 推理评测 :reasoning(abstract)

AI总结 Hyperphantasia是一个评估多模态大语言模型心理可视化能力的合成基准,通过四个精心设计的谜题揭示了人类与当前模型之间的性能差距,并探索了强化学习在提升视觉模拟能力上的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏