arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-11-04 至 2025-11-04 共收录 10 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 10 篇

2511.01203 2025-11-04 cs.LG cs.CL 86%

FEval-TTC: Fair Evaluation Protocol for Test-Time Compute

Pavel Rumiantsev, Soumyasundar Pal, Yingxue Zhang, Mark Coates

机构 * McGill University(麦吉尔大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 测试时计算 :test-time compute(title,abstract);reasoning(abstract);CoT(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05362 2025-11-04 cs.CL cs.AI cs.LG 85%

On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study

Riccardo Alberghi, Elizaveta Demyanenko, Luca Biggio, Luca Saglietti

专题命中 测试时计算 :reasoning(title,abstract);test-time compute(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00086 2025-11-04 cs.LG cs.AI cs.CL 78%

Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph

Fali Wang, Jihai Chen, Shuhua Yang, Runxue Bao, Tianxiang Zhao, Zhiwei Zhang, Xianfeng Tang, Hui Liu, Qi He, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) University of Pittsburgh(匹兹堡大学) Amazon(亚马逊) Microsoft(微软)

专题命中 测试时计算 :test-time compute(title);分类 cs.CL、cs.AI、cs.LG

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21590 2025-11-04 cs.CL cs.LG 62%

Representation Consistency for Accurate and Coherent LLM Answer Aggregation

Junqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante, Antonio Rago, Saumitra Mishra, Francesca Toni

机构 * Imperial College London(帝国理工学院伦敦分校) J.P. Morgan AI Research(摩根大通人工智能研究) King’s College London(国王学院伦敦分校)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.LG

Comments Accepted at NeurIPS 2025. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01093 2025-11-04 cs.LG cs.AI 62%

Continual Learning, Not Training: Online Adaptation For Agents

Aman Jaglan, Jarrod Barnes

机构 * Arc Intelligence

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01082 2025-11-04 cs.CV cs.AI cs.LG 62%

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi

机构 * of Computer Science, University of Southern California, Los Angeles, CA, USA Computer Engineering, University of Southern California, Los Angeles, CA, USA

专题命中 测试时计算 :test-time compute(abstract);分类 cs.AI、cs.LG

Comments Accepted to IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00711 2025-11-04 cs.LG cs.AI 62%

TRISKELION-1: Unified Descriptive-Predictive-Generative AI

Nardeep Kumar, Arun Kanwar

机构 * Independent Machine Learning Researcher(独立研究者)

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 18 figures, submitted to arXiv (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02095 2025-11-04 cs.CV cs.LG 57%

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

Hyojin Bahng, Caroline Chan, Fredo Durand, Phillip Isola

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 测试时计算 :verifier(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01423 2025-11-04 cs.SE 50%

LLM-Assisted Tool for Joint Generation of Formulas and Functions in Rule-Based Verification of Map Transformations

Ruidi He, Yu Zhang, Meng Zhang, Andreas Rausch

专题命中 测试时计算 :verifier(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01208 2025-11-04 cs.IR 50%

Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking

Jerry Huang, Siddarth Madala, Cheng Niu, Julia Hockenmaier, Tong Zhang

专题命中 测试时计算 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏