arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-11-06 至 2025-11-06 共收录 6 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 6 篇

2505.23845 2025-11-06 cs.CL 85%

Read Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs

Jakub Podolak, Rajeev Verma

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 测试时计算 :reasoning(title,abstract);chain-of-thought(abstract);test-time compute(abstract);分类 cs.CL

Comments Presented at UncertaiNLP Workshop at EMNLP 2025 https://aclanthology.org/2025.uncertainlp-main.21.pdf

Journal ref UncertaiNLP Workshop at Empirical Methods in Natural Language Processing 2025 (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00079 2025-11-06 cs.CL cs.AI 81%

PhysicsEval: Inference-Time Techniques to Improve the Reasoning Proficiency of Large Language Models on Physics Problems

Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan

机构 * Systems and Software Lab (SSL) Department of Computer Science and Engineering(系统与软件实验室(SSL)计算机科学与工程系)

专题命中 测试时计算 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted in Findings of the Association for Computational Linguistics: IJCNLP-AACL 2025, 23 pages, 4 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04825 2025-11-06 cs.GR cs.AI cs.CV cs.LG 62%

Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off

Seungyong Lee, Jeong-gi Kwak

机构 * NXN Labs(NXN实验室) University of British Columbia(不列颠哥伦比亚大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Accepted to SIGGRAPH Asia 2025, project page: https://nxnai.github.io/Voost/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03471 2025-11-06 cs.AI cs.HC 57%

Towards Scalable Web Accessibility Audit with MLLMs as Copilots

Ming Gu, Ziwei Wang, Sicen Lai, Zirui Gao, Sheng Zhou, Jiajun Bu

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

Comments 15 pages. Accepted by AAAI 2026 AISI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19248 2025-11-06 cs.LG 57%

Inference-Time Reward Hacking in Large Language Models

Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, Flavio du Pin Calmon

机构 * Harvard University(哈佛大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025 (Spotlight Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18018 2025-11-06 cs.CL 57%

Verdict: A Library for Scaling Judge-Time Compute

Nimit Kalra, Leonard Tang

机构 * Haize Labs(Haize实验室)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏