arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-01-28 至 2026-01-28 共收录 13 信号源:cs.CL, cs.AI, cs.LG

1. 复杂问题求解 13 篇

2510.20691 2026-01-28 cs.AI 85%

Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs

计划后再检索:强化学习引导的知识图谱复杂推理

Yanlin Song, Ben Liu, Víctor Gutiérrez-Basulto, Zhiwei Hu, Qianqian Xie, Min Peng, Sophia Ananiadou, Jeff Z. Pan

机构 * School of Computer Science(计算机科学学院) Wuhan University(武汉大学) School of Computer Science and Informatics(计算机科学与信息学院) Cardiff University(卡迪夫大学) College of Information Science and Engineering(信息科学与工程学院) Shanxi Agricultural University(山西农业大学) School of Artificial Intelligence(人工智能学院) Center for Language and Information Research(语言与信息研究中心) University of Manchester(曼彻斯特大学) ILCC, School of Informatics(信息学院) University of Edinburgh(爱丁堡大学)

专题命中 复杂问题求解 :reasoning(title,abstract);chain-of-thought(abstract);planning(abstract);分类 cs.AI

AI总结 Graph-RFT通过强化学习引导的知识图谱复杂推理框架,实现自主规划与适应性检索调度,解决KGQA中的冷启动问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19588 2026-01-28 cs.LG cs.AI 84%

From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation

从原子到链:基于分歧引导的推理课程用于无标签LLM领域适应

Yongqi Wang, Xiaofeng Ji, Jie Wang, Qingbin Li, Xiao Xiong, Zheming Yang, Jian Xu, Minghui Qiu, Xinxiao Wu

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science and Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) ByteDance China(字节跳动中国)

专题命中 复杂问题求解 :reasoning(title,abstract);CoT(abstract);分类 cs.AI、cs.LG

AI总结 本文提出DGRC方法,通过分歧引导构建从原子知识到推理链的学习课程,提升无标签LLM在医疗和法律领域适应性能。

Comments Code: https://github.com/bytedance/DGRC

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19290 2026-01-28 cs.CL 79%

MetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning

MetaGen: 多智能体LLM推理中的自演化角色与拓扑

Yimeng Wang, Jiaxing Zhao, Hongbin Xie, Hexing Ma, Yuzhen Lei, Shuangxue Liu, Xuan Song, Zichen Zhang, Haoran Zhang

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) School of Urban Planning and Design, Peking University(北京大学城市规划与设计学院)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL

AI总结 MetaGen通过自演化角色和拓扑提升多智能体LLM推理的准确性和效率

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06201 2026-01-28 cs.LG 79%

K2-V2: A 360-Open, Reasoning-Enhanced LLM

K2-V2:一种360开放、推理增强的LLM

K2 Team, Zhengzhong Liu, Liping Tang, Linghao Jin, Haonan Li, Nikhil Ranjan, Desai Fan, Shaurya Rohatgi, Richard Fan, Omkar Pangarkar, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Seungwook Han, Bowen Tan, Gurpreet Gosal, Xudong Han, Varad Pimpalkhute, Shibo Hao, Ming Shan Hee, Joel Hestness, Haolong Jia, Liqun Ma, Aaryamonvikram Singh, Daria Soboleva, Natalia Vassilieva, Renxi Wang, Yingquan Wu, Yuekai Sun, Taylor Killian, Alexander Moreno, John Maggs, Hector Ren, Guowei He, Hongyi Wang, Xuezhe Ma, Yuqi Wang, Mikhail Yurochkin, Eric P. Xing

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.LG

AI总结 K2-V2是一种通过注入领域知识、推理、长上下文和工具使用能力,提升复杂推理任务性能的360开放LLM,具备强大的推理能力和开源训练资源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19686 2026-01-28 cs.CV 78%

Video-KTR: Reinforcing Video Reasoning via Key Token Attribution

Video-KTR: 通过关键令牌 attribution 增强视频推理

Ziyue Wang, Sheng Jin, Zhongrong Zuo, Jiawei Wu, Han Qiu, Qi She, Hao Zhang, Xudong Jiang

机构 * ByteDance(字节跳动) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) National University of Singapore(国立新加坡大学) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 复杂问题求解 :reasoning(title,abstract)

AI总结 Video-KTR通过结合三种attribution信号,提升视频推理的准确性和可解释性,实现状态-of-the-art性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19337 2026-01-28 cs.AI cs.LG cs.SE 62%

SETA: Statistical Fault Attribution for Compound AI Systems

SETA:复合人工智能系统的统计故障归因

Sayak Chowdhury, Meenakshi D'Souza

机构 * International Institute of Information Technology Bangalore(国际信息科技学院班加罗尔)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 SETA提出了一种模块化鲁棒性测试框架,用于分析复合人工智能系统的故障归因,通过组件级分析和错误传播推理,实现细粒度的鲁棒性评估。

Comments Accepted to CAIN 2026 co-hosted with ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19871 2026-01-28 cs.CL 57%

Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection

反思性翻译:通过结构化自我反思改进低资源机器翻译

Nicholas Cheng

机构 * Independent Researcher(独立研究者)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL

AI总结 本文提出反思性翻译框架,通过结构化自我反思提升低资源语言的机器翻译质量,实验表明其在BLEU和COMET评分上均有显著提升。

Comments 12 pages, 3 figures, 6 tables. Accepted to the NeurIPS 2025 Workshop on Multilingual Representation Learning (Mexico City) and the AAAI 2025 Workshop on Language Models for Under-Resourced Communities (LM4UC). Code and data available at: https://github.com/Nickcheng123/reflective-translation-mt

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19053 2026-01-28 cs.HC cs.AI 57%

From Answer Givers to Design Mentors: Guiding LLMs with the Cognitive Apprenticeship Model

从答案提供者到设计导师:通过认知 apprenticeship 模型引导 LLMs

Yongsu Ahn, Lejun R Liao, Benjamin Bach, Nam Wook Kim

机构 * Boston College(波士顿学院) Inria(法国国家信息与自动化技术研究院) University of Edinburgh(爱丁堡大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

AI总结 本文通过认知 apprenticeship 模型引导 LLMs 作为设计导师,提升设计推理和反思反馈质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13887 2026-01-28 cs.AI 57%

Human Simulation Computation: A Human-Inspired Framework for Adaptive AI Systems

人类模拟计算:一种受人类启发的自适应人工智能系统框架

Hong Su

机构 * School of Computer Science, Chengdu University of Information Technology(计算机科学学院,成都信息科技学院)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

AI总结 本文提出人类模拟计算框架,通过结合人类思维策略与行动反馈,提升AI在动态环境中的适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23581 2026-01-28 cs.LG 57%

GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning

GraphRAG-R1: 基于过程约束强化学习的图检索增强生成

Chuanyue Yu, Kuo Zhao, Yuhan Li, Heng Chang, Mingjian Feng, Xiangzhe Jiang, Yufei Sun, Jia Li, Yuzhi Zhang, Jianxin Li, Ziwei Zhang

机构 * Nankai University(南开大学) Huawei Technologies Ltd(华为技术有限公司) Beihang University(北航)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.LG

AI总结 GraphRAG-R1通过过程约束强化学习提升LLM多跳推理能力,结合改进的GRPO方法和两种新型奖励函数,有效解决复杂问题。

Comments Accepted by the Web Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18924 2026-01-28 cs.AI 57%

RIFT: Reordered Instruction Following Testbed To Evaluate Instruction Following in Singular Multistep Prompt Structures

RIFT:重新排列指令跟随测试床以评估单步提示结构中的指令跟随

Andrew Jaffe, Noah Reicin, Jinho D. Choi

机构 * Emory University(埃默里大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

AI总结 RIFT通过测试不同提示结构下的LLM表现,揭示了指令跟随对位置连续性的强依赖,指出当前架构将指令跟随视为顺序模式而非推理能力。

Comments 13 pages, 5 figures, submitted to ACL ARR

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18847 2026-01-28 cs.SE cs.AI 57%

MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution

MulVul: 通过跨模型提示进化实现检索增强的多智能体代码漏洞检测

Zihan Wu, Jie Xu, Yun Peng, Chun Yong Chong, Xiaohua Jia

专题命中 复杂问题求解 :self-correction(abstract);分类 cs.AI

AI总结 MulVul通过跨模型提示进化实现多智能体代码漏洞检测,采用由粗到细的策略,利用检索工具提升漏洞识别精度,达到34.79%的Macro-F1成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01255 2026-01-28 cs.SE 50%

TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language Models

TestWeaver: 基于执行意识和反馈驱动的回归测试生成方法

Cuong Chi Le, Cuong Duc Van, Tung Duy Vu, Thai Minh Pham Vu, Hoang Nhat Phan, Huy Nhat Phan, Tien N. Nguyen

专题命中 复杂问题求解 :reasoning(abstract)

AI总结 TestWeaver通过整合轻量级程序分析和针对性执行上下文,提升LLM在回归测试生成中的覆盖率和效率。

Comments Accepted in ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏