arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-01-22 至 2026-01-22 共收录 3 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 3 篇

2601.14290 2026-01-22 cs.CL 83%

Project Aletheia: Verifier-Guided Distillation of Backtracking for Small Language Models

项目Aletheia:指导验证的回溯蒸馏

Aradhya Dixit, Tianxi Liang, Jai Telang

机构 * Wake Technical Community College(韦克技术社区学院) Cornell University(康奈尔大学) Algoverse

专题命中 测试时计算 :verifier(title,abstract);reasoning(abstract);分类 cs.CL

AI总结 项目Aletheia通过指导验证的蒸馏方法,使小型语言模型能够通过回溯和冲突检测提升其在约束满足问题上的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08864 2026-01-22 cs.LG cs.AI cs.CR 62%

The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses

好坏与丑:水印、可转移攻击和对抗防御的元分析

Grzegorz Głuch, Berkant Turan, Sai Ganesh Nagarajan, Sebastian Pokutta

专题命中 测试时计算 :verifier(abstract);分类 cs.AI、cs.LG

AI总结 本文通过元分析探讨了水印、可转移攻击和对抗防御之间的权衡,证明了三者中至少存在其一,并利用全同态加密构建了可转移攻击的实例。

Comments 47 pages, 3 figures, 4 tables, preliminary version published in ICML 2024 (Workshop on Theoretical Foundations of Foundation Models) and , see https://openreview.net/pdf?id=WMaFRiggwV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09345 2026-01-22 cs.CL 57%

Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time Scaling

Seer Self-Consistency:面向自适应测试时缩放的先进预算估计

Shiyu Ji, Yixuan Wang, Yijun Liu, Qingfu Zhu, Wanxiang Che

机构 * Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology, China(社会计算与交互机器人研究中心,哈尔滨工业大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

AI总结 SeerSC通过整合系统1和系统2推理,提升自适应测试时缩放的令牌效率和延迟,实现47%的令牌消耗减少和43%的推理延迟降低。

详情

展开后加载摘要…

URL PDF HTML 收藏