arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-14 至 2026-01-14 共收录 188 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 11 篇

2601.08500 2026-01-14 cs.CL 91%

It's All About the Confidence: An Unsupervised Approach for Multilingual Historical Entity Linking using Large Language Models

信心才是关键:一种用于多语言历史实体链接的无监督方法,使用大语言模型

Cristian Santini, Marieke Van Erp, Mehwish Alam

机构 * Department of Humanities, University of Macerata(马切拉塔大学人文学科系) KNAW Humanities Cluster, DHLab(荷兰人文集群、DHLab) INFRES Department, Télécom Paris(巴黎电信学院INFRES部门)

专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract);LLM(abstract);small language model(abstract)

AI总结 本文提出MHEL-LLaMo,一种结合小型语言模型和大语言模型的无监督方法,用于多语言历史实体链接,通过置信度评分降低计算成本并提升准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08405 2026-01-14 cs.RO cs.SY eess.SY 88%

Large Language Models to Enhance Multi-task Drone Operations in Simulated Environments

大语言模型提升模拟环境中的多任务无人机操作

Yizhan Feng, Hichem Snoussi, Jing Teng, Abel Cherouat, Tian Wang

专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract)

AI总结 本文提出利用大语言模型提升模拟环境中的无人机多任务操作效率,通过微调CodeT5模型实现自然语言到可执行代码的自动化翻译。

Comments 1st International Conference on Drones and Unmanned Systems (DAUS' 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08292 2026-01-14 cs.CV 88%

KidVis: Do Multimodal Large Language Models Possess the Visual Perceptual Capabilities of a 6-Year-Old?

KidVis: 多模态大语言模型是否具备六岁儿童的视觉感知能力?

Xianfeng Wang, Kaiwei Zhang, Qi Jia, Zijian Chen, Guangtao Zhai, Xiongkuo Min

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract)

AI总结 KidVis研究通过对比人类儿童与多模态大语言模型在视觉能力上的表现,揭示了当前模型在基础视觉感知上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08816 2026-01-14 cs.CL econ.EM 86%

Measuring the Quality of Answers in Political Q&As with Large Language Models

利用大语言模型评估政治问答中的答案质量

R. Michael Alvarez, Jacob Morrier

机构 * Division of the Humanities and Social Sciences(人文与社会科学系) California Institute of Technology(加州理工学院)

专题命中 预训练与数据 :language model(title,abstract);large language model(title);分类 cs.CL

AI总结 本文提出利用大语言模型评估政治问答答案质量的方法,通过语义相关性衡量答案的相关性和深度,并发现答案质量与提问议员政党存在相关性。

Journal ref Polit. Anal. 34 (2026) 78-95

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08517 2026-01-14 cs.CV 85%

Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models

闭环大语言模型在视觉模型中发现非标准通道先验

Tolgay Atinc Uzun, Dmitry Ignatov, Radu Timofte

机构 * Computer Vision Lab, CAIDAS \& IFI, University of W\"urzburg, Germany

专题命中 预训练与数据 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出利用大语言模型进行视觉模型通道配置优化,通过生成大量架构数据训练LLM学习通道配置与性能关系,实验表明该方法在CIFAR-100上显著提升准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02867 2026-01-14 cs.CL 83%

Training Language Models with homotokens Leads to Delayed Overfitting

通过homotokens训练语言模型导致过拟合延迟

Adrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai Dascalu

机构 * Dalle Molle Institute for Artificial Intelligence (IDSIA)(达勒莫莱人工智能研究所) National University of Science and Technology POLITEHNICA Bucharest(科学与技术国家大学)

专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract);分类 cs.CL

AI总结 通过homotokens训练语言模型可延迟过拟合并提升泛化能力,方法通过辅助编码器和注意力机制实现tokenization不变性。

Comments 8 pages, 6 figures, 3 Appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08773 2026-01-14 cs.SE cs.AI 79%

Reliable Graph-RAG for Codebases: AST-Derived Graphs vs LLM-Extracted Knowledge Graphs

可靠的代码库图-RAG:基于AST的图与LLM提取的知识图谱

Manideep Reddy Chinthareddy

机构 * Software Engineer, Centerville, USA(美国辛斯维尔软件工程师)

专题命中 预训练与数据 :LLM(title,abstract);分类 cs.AI

AI总结 本文比较了基于AST的图谱和LLM提取的知识图谱在代码库中的检索性能,发现确定性AST图谱在索引成本和多跳推理方面更具优势。

Comments 46 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08472 2026-01-14 cs.CL cs.AI 79%

sui-1: Grounded and Verifiable Long-Form Summarization

sui-1:可验证的长文本摘要

Benedikt Droste, Jan Philipp Harries, Maximilian Idahl, Björn Plüster

机构 * ellamind

专题命中 预训练与数据 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 sui-1通过生成带引用的摘要,解决了大型语言模型摘要不可验证的问题,展示了任务特定训练在引用支持摘要中的优越性。

Comments 13 pages, 4 figures, model weights at https://huggingface.co/ellamind/sui-1-24b

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08699 2026-01-14 cs.CL 70%

RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis

RAGShaper: 通过自动化数据合成激发复杂的代理RAG技能

Zhengwei Tao, Bo Li, Jialong Wu, Guochen Yan, Huanyao Zhang, Jiahao Xu, Haitao Mi, Wentao Zhang

机构 * Peking University(北京大学) Tencent AI Lab(腾讯人工智能实验室)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 RAGShaper通过自动化数据合成构建高质量RAG任务和代理轨迹,提升模型在复杂检索任务中的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08316 2026-01-14 cs.LG cs.CV stat.ML 70%

Deep Exploration of Epoch-wise Double Descent in Noisy Data: Signal Separation, Large Activation, and Benign Overfitting

深度探索噪声数据中的按epoch双下降现象:信号分离、大激活和良性过拟合

Tomoki Kubo, Ryuken Uda, Yusuke Iida

机构 * Niigata University(Niigata大学)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本研究通过分析噪声数据中的深度双下降现象,揭示了良性过拟合、信号分离和大激活等关键现象之间的联系。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06010 2026-01-14 cs.CV cs.MM 67%

Latent Reconstruction from Generated Data for Multimodal Misinformation Detection

从生成数据中进行潜在重建用于多模态虚假信息检测

Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis

机构 * Information Technology Institute, Centre for Research & Technology, Hellas(信息科技研究所,研究中心,希腊) Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,亚里士多德大学)

专题命中 预训练与数据 :language model(abstract);prompting(abstract)

AI总结 本研究提出MisCaption This!框架和LAMAR网络,通过生成高保真度的误标数据和潜在重建技术,提升多模态虚假信息检测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 指令微调 8 篇

2601.08141 2026-01-14 cs.CL cs.AI cs.LG 90%

Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training

Qalb:面向2.3亿使用者的最先进的乌尔都语大型语言模型,采用系统性持续预训练

Muhammad Taimoor Hassan, Jawad Ahmed, Muhammad Awais

机构 * Auburn University, USA(美国阿伯杜大学) BHT Berlin, Germany(柏林BHT学院) BTU Cottbus, Germany(库滕堡工业大学)

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);foundation model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Qalb通过持续预训练和监督微调,为乌尔都语构建了最先进的大型语言模型,显著提升了在多种任务上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07198 2026-01-14 cs.LG 83%

Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning

协同优于差异:一种基于分区的多领域LLM微调方法

Hua Ye, Siyuan Chen, Haoliang Zhang, Weihao Luo, Yanbin Li, Xuan Zhang

机构 * Nanjing University(南京大学) Airon Technology CO., LTD(艾瑞森技术有限公司) University of Bristol(布里斯托大学) The University of Oklahoma(俄克拉荷马大学) Donghua University(东华大学) Beijing University of Posts and Telecommunications(北京邮电大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 指令微调 :LLM(title);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出一种基于分区的多领域LLM微调方法,通过平衡领域差异与协同效应,有效减少领域间干扰,提升多领域适应性能。

Comments 20 pages, 5 figures, 21 tables. Accepted at NeurIPS 2025. Corresponding author: Xuan Zhang (xuanzhang2199@gmail.com)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16913 2026-01-14 cs.SE 75%

FGIT: Fault-Guided Fine-Tuning for Code Generation

FGIT:基于故障的微调用于代码生成

Lishui Fan, Zhongxin Liu, Haoye Wang, Lingfeng Bao, Xin Xia, Shanping Li

专题命中 指令微调 :large language model(abstract);language model(abstract);SFT(abstract)

AI总结 FGIT通过提取正确与错误代码差异并动态加权提升代码生成准确性,实现性能提升。

Comments 13 pages, accepted by ASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08198 2026-01-14 cs.CL cs.LG 73%

Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs

三元组优于对:迈向稳定且有效的自play微调方法用于大语言模型

Yibo Wang, Hai-Long Sun, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Lijun Zhang

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室) School of Artificial Intelligence, Nanjing University(人工智能学院) Alibaba International Digital Commerce(阿里巴巴国际数字商务) Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔))

专题命中 指令微调 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 T-SPIN通过引入历史优势和熵约束,提升自play微调在稀缺标注数据下的稳定性和性能

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01124 2026-01-14 cs.IR 71%

TransFR: Transferable Federated Recommendation with Adapter Tuning on Pre-trained Language Models

TransFR: 基于预训练语言模型的可迁移联邦推荐

Honglei Zhang, Zhiwei Li, Haoxuan Li, Xin Zhou, Jie Zhang, Yidong Li

专题命中 指令微调 :language model(title)

AI总结 TransFR通过结合预训练语言模型的通用能力和本地数据微调的个性化能力,解决联邦推荐中的可迁移性、冷启动和隐私问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10352 2026-01-14 cs.CL 70%

Cross-Prompt Encoder for Low-Performing Languages

跨提示编码器用于表现不佳的语言

Beso Mikaberidze, Teimuraz Saghinadze, Simon Ostermann, Philipp Muller

机构 * Muskhelishvili Institute of Computational Mathematics, GTU (MICM)(穆斯赫利什维利计算数学研究所(MICM)) Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI)(德国人工智能研究中心(DFKI)) Center for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心(CERTAIN)) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 指令微调 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出跨提示编码器(XPE)用于提升表现不佳语言的性能,并结合双软提示机制增强多语言适应能力。

Comments Accepted at Findings of IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04954 2026-01-14 cs.LG cs.AI 62%

Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following

精确性胜过多样性:高精度奖励能泛化到稳健的指令跟随

Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou, Yuxian Wang, Haonan Song, Wu Ning, Dandan Tu, Qixun Zhang, Bibo Cai, Yuxiang He, Ting Liu

机构 * Harbin Institute of Technology, SCIR(哈尔滨工业大学,SCIR) Peking University(北京大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 指令微调 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 本研究发现高精度奖励优于多样化的约束混合,提出数据导向的优化策略,提升指令跟随性能并减少训练时间。

Comments Under review, 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08486 2026-01-14 physics.comp-ph 50%

Multi-Task Fine-Tuning Enables Robust Out-of-Distribution Generalization in Atomistic Models

多任务微调使原子模型在分布外泛化中更加稳健

Chengqian Zhang, Duo Zhang, Anyang Peng, Mingyu Guo, Yuzhi Zhang, Lei Wang, Guolin Ke, Linfeng Zhang, Tiejun Li, Han Wang

专题命中 指令微调 :pretraining(abstract)

AI总结 多任务微调通过联合优化性质预测与物理基础力场目标,提升了原子模型在分布外情况下的泛化能力,优于传统微调和任务特定模型。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 后训练与偏好优化 8 篇

2511.21005 2026-01-14 cs.AI cs.IR 87%

ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning

ICPO:内在置信度驱动的群体相对偏好优化用于高效强化学习

Jinpeng Wang, Chao Li, Ting Ye, Mengyuan Zhang, Wei Liu, Jian Luan

机构 * MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 ICPO通过内在置信度驱动的群体相对偏好优化,提升大型语言模型的推理能力,有效解决粗粒度奖励和奖励噪声问题,增强高质量响应的相对优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14381 2026-01-14 cs.LG cs.AI cs.CL cs.CR 87%

Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers

我的优化提示是否被 compromised?探索基于 LLM 的优化器的漏洞

Andrew Zhao, Reshmi Ghosh, Vitor Carvalho, Emily Lawton, Keegan Hines, Gao Huang, Jack W. Stokes

机构 * Tsinghua University(清华大学) Microsoft(微软)

专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现基于 LLM 的提示优化存在重大安全漏洞,提出假奖励攻击及轻量级防御措施,揭示优化流程为重要攻击目标。

Comments Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22462 2026-01-14 cs.AI 81%

Learning "Partner-Aware" Collaborators in Multi-Party Collaboration

在多方协作中学习“伙伴感知”的协作者

Abhijnan Nath, Nikhil Krishnaswamy

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);RLHF(abstract)

AI总结 本文提出ICR算法,通过学习伙伴感知的协作者,提升多任务协作中的共同基础一致性与解决方案多样性。

Comments Fixed typographic errors in the previous manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08777 2026-01-14 cs.LG cs.AI cs.CL cs.GT 80%

Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling

渐近通用对齐:通过测试时间缩放实现的新对齐框架

Yang Cai, Weiqiang Zheng

机构 * Yale University(耶鲁大学)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种基于测试时间缩放的通用对齐框架,通过多玩家对齐游戏实现最优鲁棒性,解决了现有方法在输出多样性上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02618 2026-01-14 cs.CL 77%

Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation

通过交互蒸馏缓解判别奖励建模中的注意力黑客问题

Jianxiang Zang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);RLHF(abstract);分类 cs.CL

AI总结 本文提出交互蒸馏方法,通过优化注意力机制提升判别奖励建模的稳定性与通用性,缓解注意力黑客问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26167 2026-01-14 cs.AI cs.CL 73%

ToolRM: Towards Agentic Tool-Use Reward Modeling

ToolRM: 向代理工具使用奖励建模迈进

Renhao Li, Jianhong Tu, Yang Su, Yantao Liu, Fei Huang, Hamid Alinejad-Rokny, Derek F. Wong, Junyang Lin, Min Yang

机构 * University of Macau(澳门大学) Qwen Team, Alibaba Inc.(阿里云Qwen团队) UNSW Sydney(新南威尔士大学悉尼分校) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 ToolRM通过构建高质量偏好数据集和基准,提升代理AI在工具调用任务中的表现和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08237 2026-01-14 cs.AI 70%

The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination

奖励工程的终结:大型语言模型如何重新定义多智能体协调

Haoran Su, Yandong Sun, Congjia Yu

机构 * New York University(纽约大学) Lerna AI

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文探讨了大型语言模型如何通过语义奖励规范和动态适应替代传统奖励工程,重新定义多智能体协调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08421 2026-01-14 cs.LG 57%

Coverage Improvement and Fast Convergence of On-policy Preference Learning

策略学习中的覆盖改进与快速收敛

Juno Kim, Jihun Yun, Jason D. Lee, Kwang-Sung Jun

机构 * UC Berkeley(伯克利大学) KRAFTON(KRAFTON公司) University of Arizona(亚利桑那大学)

专题命中 后训练与偏好优化 :language model(abstract);分类 cs.LG

AI总结 本文提出覆盖改进原理,通过分析在线策略学习中的覆盖变化,证明其在批量大小足够时能实现快速收敛,并设计混合采样器和奖励蒸馏方案,提升语言模型对齐性能。

Comments 46 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长上下文与记忆 11 篇

2601.08829 2026-01-14 cs.CL cs.AI 86%

Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System

在Elo排名评审系统中建模LLM代理评审动态

Hsiang-Wei Huang, Junbin Lu, Kuang-Ming Chen, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学)

专题命中 长上下文与记忆 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了在Elo排名评审系统中,通过整合Elo评分和评审者记忆提升领域主席决策准确性的LLM代理评审动态。

Comments In submission. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08739 2026-01-14 cs.CL 85%

PrivGemo: Privacy-Preserving Dual-Tower Graph Retrieval for Empowering LLM Reasoning with Memory Augmentation

PrivGemo: 保护隐私的双塔图检索用于增强LLM推理的记忆增强

Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu, Xin Yuan, Liming Zhu, Wenjie Zhang

机构 * University of New South Wales(新南威尔士大学) Data61, CSIRO(Data61,CSIRO)

专题命中 长上下文与记忆 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 PrivGemo通过双塔设计和隐私保护机制,实现隐私保护的图检索增强推理,提升LLM在知识密集型任务中的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08215 2026-01-14 cs.CL cs.LG 81%

Towards Principled Design of Mixture-of-Experts Language Models under Memory and Inference Constraints

面向内存和推理约束下混合专家语言模型的原理性设计

Seng Pei Liew, Kenta Shinzato, Yuyang Dong

机构 * SB Intuitions

专题命中 长上下文与记忆 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出基于内存和推理约束的混合专家语言模型设计原则,通过系统研究揭示性能主要由总参数和专家稀疏性决定,并提出优化方法指导架构设计。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏