arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 2047 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 2047 篇

2402.10151 2024-02-16 cs.CL 57%

ControlLM: Crafting Diverse Personalities for Language Models

Yixuan Weng, Shizhu He, Kang Liu, Shengping Liu, Jun Zhao

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14262 2023-12-25 cs.SE cs.AI 57%

Exploring the intersection of Generative AI and Software Development

Filipe Calegario, Vanilson Burégio, Francisco Erivaldo, Daniel Moraes Costa Andrade, Kailane Felix, Nathalia Barbosa, Pedro Lucas da Silva Lucena, César França

专题命中 测试时计算 :chain-of-thought(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02541 2023-11-29 cs.CR cs.AI 57%

PCPT and ACPT: Copyright Protection and Traceability Scheme for DNN Models

Xuefeng Fan, Dahao Fu, Hangyu Gui, Xinpeng Zhang, Xiaoyi Zhou

专题命中 测试时计算 :verifier(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04495 2023-11-09 cs.CL 57%

Multi-label and Multi-target Sampling of Machine Annotation for Computational Stance Detection

Zhengyuan Liu, Hai Leong Chieu, Nancy F. Chen

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

Comments Findings of EMNLP 2023. arXiv admin note: text overlap with arXiv:2305.19845

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08303 2023-09-18 cs.CL 57%

Self-Consistent Narrative Prompts on Abductive Natural Language Inference

Chunkit Chan, Xin Liu, Tsz Ho Chan, Jiayang Cheng, Yangqiu Song, Ginny Wong, Simon See

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2023 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07645 2023-08-21 cs.CL 57%

Steering Language Generation: Harnessing Contrastive Expert Guidance and Negative Prompting for Coherent and Diverse Synthetic Data Generation

Charles O'Neill, Yuan-Sen Ting, Ioana Ciuca, Jack Miller, Thang Bui

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00108 2023-08-21 cs.SE cs.LG 57%

Better patching using LLM prompting, via Self-Consistency

Toufique Ahmed, Premkumar Devanbu

专题命中 测试时计算 :CoT(abstract);分类 cs.LG

Comments Accepted at ASE-NIER (2023) track

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08204 2023-08-17 cs.CL 57%

MoCoSA: Momentum Contrast for Knowledge Graph Completion with Structure-Augmented Pre-trained Language Models

Jiabang He, Liu Jia, Lei Wang, Xiyao Li, Xing Xu

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02072 2022-11-02 cs.LG cs.IT math.IT stat.ML 57%

Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning

Dilip Arumugam, Benjamin Van Roy

专题命中 测试时计算 :planning(abstract);分类 cs.LG

Comments Accepted to Neural Information Processing Systems (NeurIPS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01975 2022-09-07 cs.CL 57%

Selective Annotation Makes Language Models Better Few-Shot Learners

Hongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10615 2022-05-23 cs.CL cs.LO 57%

Generalized Quantifiers as a Source of Error in Multilingual NLU Benchmarks

Ruixiang Cui, Daniel Hershcovich, Anders Søgaard

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL

Comments To appear at NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04061 2022-05-10 cs.CV cs.AI 57%

Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering

Min Peng, Chongyang Wang, Yuan Gao, Yu Shi, Xiang-Dong Zhou

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

Comments Accepted by IJCAI 2022. arXiv admin note: text overlap with arXiv:2109.04735

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.09093 2021-11-18 cs.AI 57%

The Faulty GPS Problem: Shortest Time Paths in Networks with Unreliable Directions

Steve Alpern

专题命中 测试时计算 :planning(abstract);分类 cs.AI

Comments 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.07495 2021-02-16 cs.GT cs.AI cs.NE 57%

ScrofaZero: Mastering Trick-taking Poker Game Gongzhu by Deep Reinforcement Learning

Naichen Shi, Ruichen Li, Sun Youran

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

Comments The very first versoin. Will be improved in the future

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.04471 2018-11-13 cs.LG stat.ML 57%

Thompson Sampling for Pursuit-Evasion Problems

Zhen Li, Nicholas J. Meyer, Eric B. Laber, Robert Brigantic

专题命中 测试时计算 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.03078 2017-05-10 cs.AI 57%

An Anthropic Argument against the Future Existence of Superintelligent Artificial Intelligence

Toby Pereira

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1302.2828 2013-02-13 cs.RO cs.AI cs.MA 57%

Multi-agent RRT*: Sampling-based Cooperative Pathfinding (Extended Abstract)

Michal Čáp, Peter Novák, Jiří Vokřínek, Michal Pěchouček

专题命中 测试时计算 :planning(abstract);分类 cs.AI

Comments To appear at AAMAS 2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1203.3528 2012-03-19 cs.AI 57%

Rollout Sampling Policy Iteration for Decentralized POMDPs

Feng Wu, Shlomo Zilberstein, Xiaoping Chen

专题命中 测试时计算 :planning(abstract);分类 cs.AI

Comments Appears in Proceedings of the Twenty-Sixth Conference on Uncertainty in Artificial Intelligence (UAI2010)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22468 2026-08-25 math.CO 新提交 56%

An Explicit 82-Queen Covering of the 163 x 163 Board and Its Asymptotic Implication

163×163棋盘的显式82个皇后覆盖及其渐近意义

Yixiang Kong

专题命中 测试时计算 :verifier(abstract,comments)

AI总结 该研究构造了163×163棋盘的显式82个皇后A型1-覆盖,确定γ(Q₁₆₃)=82,并利用其得到更优的皇后图覆盖数渐近系数,还提供了坐标与Python验证器。

Comments 6 pages; includes complete coordinates, a standard-library Python verifier, and a machine-readable certificate. AI-assisted programming and manuscript preparation are disclosed in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21884 2026-08-28 cs.SE 版本更新 50%

Loop Engineering: Building Blocks, Adoption, and Impact

循环工程:构建模块、采用情况及影响

Jai Lal Lulla, Vahram Nersesyan, Seyedmoein Mohsenimofidi, Christoph Treude, Sebastian Baltes

专题命中 测试时计算 :verifier(abstract)

AI总结 本文对新兴循环工程领域开展探索性研究,分析其构建模块、开源项目采用情况,挖掘36710个仓库发现217个存在自主智能体循环,并规划了智能体自主层级的对照研究。

Comments 10 pages, 2 tables, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25299 2026-08-27 cs.CV 新提交 50%

PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

PointRL:从可验证标注证据中学习点级视觉-语言接地

Jingyang Su, Pu Cao, Xiuze Jin, Longyue Zhang, Qing Song, Lu Yang

机构 * School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(北京邮电大学智能工程与自动化学院)

专题命中 测试时计算 :verifier(abstract)

AI总结 PointRL是一种可验证强化学习框架,通过将异构标注证据转换为指向指令并设计多维度奖励,提升了Qwen3.5-4B等模型的点级视觉-语言接地性能,在多个基准上取得显著增益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24667 2026-08-26 cs.IR 新提交 50%

EviGraph: Towards Verifiable Evidence Construction for Information-Seeking Agents

EviGraph:面向信息检索智能体的可验证证据构建

Jiashun Chen, Yirong Mao, Wenhui Que

专题命中 测试时计算 :verifier(abstract)

AI总结 EviGraph是分离搜索与证据记录的深度搜索框架,在BrowseComp等数据集上显著提升了智能体网络搜索的准确率,每轮生成token更少,验证了结构化证据记录的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22868 2026-08-25 cs.CR 新提交 50%

AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems

AgentFlow:一种以流为中心的策略语言与框架,用于保障大语言模型智能体系统的安全

Basavesh Ammanaghatta Shivakumar, Swarn Priya, Peng Gao

专题命中 测试时计算 :verifier(abstract)

AI总结 本文提出以流为中心的AgentFlow策略语言与框架,通过运行时监控与SMT验证保障LLM智能体系统安全,在多基准测试中显著降低入侵率并提升效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28147 2026-08-25 cs.CR 版本更新 50%

Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems

智能体调控蒸馏:自主多智能体系统中的推理时调控提取与利用

Yu Cui, Wuli Yang, Yirui Shi, Junhao Xia, Hui Jiang, Lei Gao, Chenfu Bao

专题命中 测试时计算 :reasoning(abstract)

AI总结 本文提出AHD框架,将AMAS推理时调控提取形式化为新安全问题,实验证实其可提取调控致IP泄露,还提出对应欺骗防御,揭示AMAS新安全威胁。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17095 2026-08-19 cs.CV 新提交 50%

Inference-Time Attention Steering for Vision-Language-Action Driving Models

面向视觉-语言-动作驾驶模型的推理时注意力引导

Darshan Nagendra Prasad, Lars Ullrich, Knut Graichen

机构 * FAU Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)

专题命中 测试时计算 :reasoning(abstract)

AI总结 该研究针对VLA驾驶模型无法在推理时重定向注意力的问题,在Qwen3-VL主干上添加预softmax注意力偏差,实验验证了其对轨迹的引导效果及层相关特性。

Comments Attention Steering, Vision-Language-Action, AutonomousDriving, Inference-Time Intervention

Journal ref European Conference on Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11076 2026-08-12 cs.CV 新提交 50%

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

基于基础模型的高效数据采样(FEEDS):用于泛癌、多示踪剂PET/CT数据集的标签高效训练策略

Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James B. Yu, Indrani Bhattacharya

机构 * Geisel School of Medicine at Dartmouth(达特茅斯盖泽尔医学院) Dartmouth Hitchcock Medical Center(达特茅斯-希区柯克医疗中心) Yale University(耶鲁大学)

专题命中 测试时计算 :planning(abstract)

AI总结 该研究提出FEEDS策略,利用视觉基础模型嵌入选择信息性和多样性未标注病例,在减少70%标注负担时达到全标注训练性能,解决了PET/CT病灶分割的标签稀缺问题。

Comments Code is publicly available on https://github.com/Image-and-Multimodal-Data-Analytics/FEEDS

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09344 2026-08-11 cs.CV 新提交 50%

Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

超越全局编辑:用于LVLMs无训练幻觉缓解的实例级解耦子空间

Ali Cheraghian, Hamidreza Dastmalchi, Hamed Barzamini, Morteza Saberi, Mojtaba Golzan, Shafin Rahman, Hossein Rahmani

机构 * Macquarie University(麦考瑞大学) York University(约克大学) Northern Illinois University(北伊利诺伊大学) University of Technology Sydney(悉尼科技大学) North South University(北南大学) Lancaster University(兰卡斯特大学) Australian National University(澳大利亚国立大学)

专题命中 测试时计算 :reasoning(abstract)

AI总结 针对LVLMs现有模型编辑技术采用单一全局子空间无法适配多样幻觉模式的问题,提出无训练的实例级解耦子空间框架,经多基准实验验证其鲁棒性、通用性与效率。

Comments BMVC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07555 2026-08-11 cs.RO 新提交 50%

You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy Improvement

你无需停留在循环中:用于机器人策略改进的智能体机器人循环

Hang Yu

专题命中 测试时计算 :verifier(abstract)

AI总结 该研究将编码智能体的软件循环架构迁移至机器人策略改进领域,提出AgenticRobotics控制平面,解决机器人工具易失效问题,实现错误提升控制等性能提升。

Comments Agentic Robotics Preview Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06065 2026-08-07 cs.CV 新提交 50%

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

下一张截图知晓:面向移动GUI智能体的门控事后蒸馏(Gated Hindsight Distillation)

Weiwei Li, Junzhuo Liu, Tong Chu, Hengfu Yu, Wen Li

专题命中 测试时计算 :reasoning(abstract)

AI总结 本文针对移动GUI智能体训练中丢失动作依据的问题,提出门控事后蒸馏方法,利用下一张截图作为特权信息,在AndroidWorld等基准上相比GRPO提升了任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05592 2026-08-07 cs.CV 新提交 50%

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

超越帧选择:用多模态大语言模型重新思考长视频理解

Ziling Huang, Shin'ichi Satoh

机构 * National Institute of Informatics(信息学研究所)

专题命中 测试时计算 :reasoning(abstract)

AI总结 针对 MLLMs 长视频理解的帧选择策略难以兼顾全局与局部信息的问题,提出 VideoRouter 模型,通过时间层次结构和验证引导路由器协调全局与局部推理,在 VideoMME 数据集上实现性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏