arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-11-06 至 2025-11-06 共收录 56 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 6 篇

2511.03471 2025-11-06 cs.AI cs.HC 57%

Towards Scalable Web Accessibility Audit with MLLMs as Copilots

Ming Gu, Ziwei Wang, Sicen Lai, Zirui Gao, Sheng Zhou, Jiajun Bu

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

Comments 15 pages. Accepted by AAAI 2026 AISI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19248 2025-11-06 cs.LG 57%

Inference-Time Reward Hacking in Large Language Models

Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, Flavio du Pin Calmon

机构 * Harvard University(哈佛大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025 (Spotlight Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18018 2025-11-06 cs.CL 57%

Verdict: A Library for Scaling Judge-Time Compute

Nimit Kalra, Leonard Tang

机构 * Haize Labs(Haize实验室)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 复杂问题求解 9 篇

2508.03159 2025-11-06 cs.LG cs.AI 90%

CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction

Jueon Park, Yein Park, Minju Song, Soyon Park, Donghyeon Lee, Seungheun Baek, Jaewoo Kang

机构 * Department of Computer Science and Engineering(计算机科学与工程系) AIGEN Sciences(AIGEN公司)

专题命中 复杂问题求解 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI、cs.LG

Comments Accepted to IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21861 2025-11-06 cs.LG cs.AI cs.CL 87%

The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems

Bentley DeVilling

专题命中 复杂问题求解 :reasoning(title,abstract);verifier(abstract);self-correction(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 2 figures. Category: cs.LG. Code and data: https://github.com/Course-Correct-Labs/mirror-loop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02834 2025-11-06 cs.AI cs.CL cs.LG 82%

Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything

Huawei Lin, Yunzhi Shi, Tong Geng, Weijie Zhao, Wei Wang, Ravender Pal Singh

机构 * Amazon(亚马逊公司) Rochester Institute of Technology(罗切斯特理工学院) University of Rochester(罗切斯特大学)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 16 pages, 7 figures, 14 tables. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02219 2025-11-06 cs.AI 79%

TabDSR: Decompose, Sanitize, and Reason for Complex Numerical Reasoning in Tabular Data

Changjiang Jiang, Fengchang Yu, Haihua Chen, Wei Lu, Jin Zeng

机构 * Wuhan University(武汉大学) University of North Texas(北德克萨斯大学) Hubei University of Economics(湖北经济学院)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Findings

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03328 2025-11-06 cs.CL cs.AI cs.CV cs.LG 67%

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

Jindong Hong, Tianjie Chen, Lingjie Luo, Chuanyang Zheng, Ting Xu, Haibao Yu, Jianing Qiu, Qianzhong Chen, Suning Huang, Yan Xu, Yong Gui, Yijun He, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学) Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03206 2025-11-06 cs.CV cs.AI cs.LG 62%

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

Kuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 16 pages

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14234 2025-11-06 cs.AI cs.MA 57%

Beyond Single Pass, Looping Through Time: KG-IRAG with Iterative Knowledge Retrieval

Ruiyi Yang, Hao Xue, Imran Razzak, Hakim Hacid, Flora D. Salim

机构 * University of New South Wales(新南威尔士大学) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Technology Innovation Institute(技术创新研究所)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03023 2025-11-06 cs.AI 57%

PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework

Sina Montazeri, Yunhe Feng, Kewei Sha

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03478 2025-11-06 cs.HC 50%

SVG Decomposition for Enhancing Large Multimodal Models Visualization Comprehension: A Study with Floor Plans

Jeongah Lee, Ali Sarvghad

专题命中 复杂问题求解 :reasoning(abstract)

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 推理评测 9 篇

2506.21448 2025-11-06 eess.AS cs.CV cs.SD 89%

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang, Qian Chen, Zhou Zhao, Wei Xue

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) Tongyi Fun Team, Alibaba Group(阿里云团队) Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract)

Comments Accepted by NeurIPS 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04201 2025-11-06 cs.CV cs.AI 79%

ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs

Ben Zhang, LuLu Yu, Lei Gao, QuanJiang Guo, Jing Liu, Hui Gao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18140 2025-11-06 cs.CL 79%

MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning

Xiaoyuan Li, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu

机构 * University of Science and Technology of China(中国科学技术大学) Alibaba Group(阿里巴巴集团) National University of Singapore(新加坡国立大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03635 2025-11-06 cs.CL cs.LG 62%

Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability

Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

Comments Accepted in AAAI CONFERENCE ON WEB AND SOCIAL MEDIA (ICWSM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16638 2025-11-06 cs.CL cs.AI 62%

Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation

Sanjana Ramprasad, Byron C. Wallace

机构 * Northeastern University(东北大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03376 2025-11-06 eess.IV cs.AI q-bio.QM 57%

Computational Imaging Meets LLMs: Zero-Shot IDH Mutation Prediction in Brain Gliomas

Syed Muqeem Mahmood, Hassan Mohy-ud-Din

机构 * School of Science and Engineering, Lahore University of Management Sciences, Lahore, Pakistan(科学与工程学院,拉合尔管理科学大学,巴基斯坦拉合尔)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 5 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02589 2025-11-06 cs.AI 57%

The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models

Claudia Herambourg, Dawid Siuda, Julia Kopczyńska, Joao R. L. Santos, Wojciech Sas, Joanna Śmietańska-Nowak

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03106 2025-11-06 cs.AI 57%

Large language models require a new form of oversight: capability-based monitoring

Katherine C. Kellogg, Bingyang Ye, Yifan Hu, Guergana K. Savova, Byron Wallace, Danielle S. Bitterman

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03178 2025-11-06 cs.CV 50%

SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention

Shreyas C. Dhake, Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque

机构 * UCL Hawkes Institute(UCL哈维斯研究所) University College London(伦敦大学学院) Dept of Medical Physics & Biomedical Engineering(医学物理与生物医学工程系) UCL(伦敦大学学院) Dept of Computer Science(计算机科学系) National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院) Division of Informatics, Imaging and Data Science(信息学、成像与数据科学 division)

专题命中 推理评测 :reasoning(abstract)

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他推理 5 篇

2508.20637 2025-11-06 cs.LG cs.AI cs.CL 82%

GDS Agent for Graph Algorithmic Reasoning

Borun Shi, Ioannis Panagiotas

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18293 2025-11-06 cs.CL cs.AI cs.CY 73%

Evaluating Large Language Models for Detecting Antisemitism

Jay Patel, Hrudayangam Mehta, Jeremy Blackburn

机构 * Binghamton University(宾夕法尼亚州立大学)

专题命中 其他推理 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13890 2025-11-06 cs.CL cs.AI 62%

A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness

Fali Wang, Jihai Chen, Shuhua Yang, Ali Al-Lawati, Linli Tang, Hui Liu, Suhang Wang

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Michigan State University(密歇根州立大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 19 figures-under review; more detailed than v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03152 2025-11-06 cs.CL cs.AI 62%

Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment

Srishti Yadav, Jasmina Gajcin, Erik Miehling, Elizabeth Daly

机构 * IBM Research, Ireland(IBM研究院(爱尔兰))

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02558 2025-11-06 cs.CL 57%

Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction

Yuerong Song, Xiaoran Liu, Ruixiao Li, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Ziwei He, Xipeng Qiu

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏