arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-26 至 2026-01-26 共收录 17 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 17 篇

2601.00364 2026-01-26 cs.CL 92%

The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining

混合语言文档在多语言大语言模型预训练中的作用

Jiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani, Pontus Stenetorp, Jianfei Yang, Yao Lu

机构 * University College London(伦敦大学学院) Nanyang Technological University(南洋理工大学) University of Waterloo(滑铁卢大学) NVIDIA(英伟达) National Institute of Informatics(信息研究院)

专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract);pretraining(title,abstract);分类 cs.CL

AI总结 研究发现平行数据对翻译性能至关重要,而语码转换数据对跨语言理解和推理影响较小。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15892 2026-01-26 cs.CL 88%

Stable-DiffCoder: Pushing the Frontier of Code Diffusion Large Language Model

Stable-DiffCoder:推动代码扩散大语言模型的前沿

Chenghao Fan, Wen Heng, Bo Li, Sichen Liu, Yuxuan Song, Jing Su, Xiaoye Qu, Kai Shen, Wei Wei

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 预训练与数据 :language model(title,abstract);large language model(title);pretraining(abstract);分类 cs.CL

AI总结 Stable-DiffCoder通过基于扩散的训练方法,在代码生成任务中超越了传统自回归模型,提升了代码建模的质量和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16618 2026-01-26 cs.CL 87%

PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs

PROST-LLM:逐步增强大语言模型中的语音到语音翻译能力

Jing Xu, Jiaqi Wang, Daxin Tan, Xiao Chen

机构 * The Chinese University of Hong Kong Huawei Artificial Intelligence Laboratory (Leibniz)(香港中文大学华为人工智能实验室(莱布尼茨))

专题命中 预训练与数据 :LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)

AI总结 PROST-LLM通过逐步增强方法提升大语言模型在语音到语音翻译中的能力,采用三任务学习和模态链方法,结合自我采样和回译生成偏好对,最终通过偏好优化提升翻译性能。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16278 2026-01-26 cs.CL cs.AI cs.LG 83%

Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification

优于分类器:利用大语言模型和合成数据进行低资源多语言分类

Branislav Pecher, Jan Cegin, Robert Belanec, Ivan Srba, Jakub Simko, Maria Bielikova

机构 * Kempelen Institute of Intelligent Technologies(克姆佩尔智能技术研究所) Faculty of Information Technology, Brno University of Technology(信息科技学院)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);instruction tuning(abstract)

AI总结 本研究利用大语言模型生成合成数据,训练更小的多语言模型,发现生成器在低资源语言中表现更优。

Comments Accepted to the Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21548 2026-01-26 cs.CL cs.AI cs.CY physics.soc-ph 82%

Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment

流畅但异域:即使区域LLM也缺乏文化契合

Dhruv Agarwal, Anya Shukla, Sunayana Sitaram, Aditya Vashistha

机构 * Cornell University(康奈尔大学) Microsoft Research(微软研究院)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);pretraining(abstract);prompting(abstract)

AI总结 研究发现,即使区域LLM也缺乏文化契合,需通过社区基础的数据和厚宽评估来构建真正主权的LLM。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16451 2026-01-26 cs.CV 82%

VISTA-PATH: An interactive foundation model for pathology image segmentation and quantitative analysis in computational pathology

VISTA-PATH: 一种交互式的基础模型用于计算病理学中病理图像分割和定量分析

Peixian Liang, Songhao Li, Shunsuke Koga, Yutong Li, Zahra Alipour, Yucheng Tang, Daguang Xu, Zhi Huang

机构 * Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学) Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学) Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学) NVIDIA Corporation(NVIDIA公司) Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)

专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)

AI总结 VISTA-PATH是一种交互式基础模型,通过整合专家反馈和多类分割,提升病理图像分割的临床应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09954 2026-01-26 cs.CV 82%

The Spatial Blindspot of Vision-Language Models

视觉-语言模型的空间盲区

Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu, Timothy Chung, Drishti Sharma, Akshata A, Kranthi Kiran, Wesley Tam, Bala Krishna S Vegesna

专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract)

AI总结 本文提出通过改进图像编码器和2D位置编码来提升视觉-语言模型的空间推理能力,以解决其在空间关系捕捉上的不足。

Comments Work done as part of the EleutherAI SOAR Program

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23407 2026-01-26 cs.LG cs.AI cs.CL 80%

Theoretical Foundations of Scaling Law in Familial Models

家族模型中的缩放定律理论基础

Huan Song, Qingfei Zhao, Ting Long, Shuyu Tian, Hongjun An, Jiawei Shao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出家族模型的缩放定律理论,通过引入粒度变量扩展传统缩放定律,证明在不牺牲计算最优性的情况下可实现多次部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13223 2026-01-26 cs.LG 79%

Data Matters Most: Auditing Social Bias in Contrastive Vision Language Models

数据最为关键:审计对比视觉语言模型中的社会偏见

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * Queen Mary University of London(伦敦女王学院) Institut Jožef Stefan(Jožef Stefan研究所)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.LG

AI总结 研究通过对比CLIP和OpenCLIP模型,发现数据来源是偏见的主要驱动因素,不同去偏策略在不同模型和数据规模下效果各异。

Comments Published at TMLR; updated version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16492 2026-01-26 cs.IR 78%

LLM-based Semantic Search for Conversational Queries in E-commerce

基于大语言模型的语义搜索用于电商会话查询

Emad Siddiqui, Venkatesh Terikuti, Xuan Lu

专题命中 预训练与数据 :LLM(title,abstract)

AI总结 本文提出基于大语言模型的语义搜索框架,通过结合领域嵌入与结构化过滤,提升电商会话查询的检索精度与召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16354 2026-01-26 cs.CR cs.AI 77%

NOIR: Privacy-Preserving Generation of Code with Open-Source LLMs

NOIR:基于开源LLM的隐私保护代码生成

Khoa Nguyen, Khiem Ton, NhatHai Phan, Issa Khalil, Khang Tran, Cristian Borcea, Ruoming Jin, Abdallah Khreishah, My T. Thai

机构 * New Jersey Institute of Technology(新泽西理工学院) Hamad Bin Khalifa University(哈利法大学) Kent State University(肯特州立大学) University of Florida(佛罗里达大学)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 NOIR通过客户端本地处理和隐私保护机制,实现对开源LLM生成代码的隐私保护,显著提升代码生成性能和安全性。

Comments To appear at Usenix Security Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00505 2026-01-26 cs.CL 77%

Zero-RAG: Towards Retrieval-Augmented Generation with Zero Redundant Knowledge

零冗余知识下的检索增强生成:Zero-RAG

Qi Luo, Xiaonan Li, Junqi Dai, Shuang Cheng, Xipeng Qiu

机构 * School of Computer Science, Fudan University(复旦大学计算机学院)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 Zero-RAG通过修剪冗余知识和优化检索流程,提升LLM在减少冗余信息下的生成性能。

Journal ref Frontiers of Computer Science (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10805 2026-01-26 cs.LG 77%

Detecting High-Stakes Interactions with Activation Probes

通过激活探针检测高风险交互

Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger, Ekdeep Singh Lubana, Dmitrii Krasheninnikov

机构 * LASR Labs(LASR实验室) University College London(伦敦大学学院) MILA Harvard University(哈佛大学) NTT Research(NTT研究所) Goodfire University of Cambridge(剑桥大学)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 通过激活探针检测高风险交互,实现高效且资源敏感的监控系统

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16541 2026-01-26 cs.CV cs.LG 57%

Semi-Supervised Hierarchical Open-Set Classification

半监督层次开放式分类

Erik Wallin, Fredrik Kahl, Lars Hammarstrand

专题命中 预训练与数据 :pretraining(abstract);分类 cs.LG

AI总结 本文提出了一种半监督层次开放式分类框架,通过子树伪标签和年龄门控机制,在有限标注数据下提升分类性能。

Comments WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00564 2026-01-26 cs.LG physics.flu-dyn 57%

Pre-Generating Multi-Difficulty PDE Data for Few-Shot Neural PDE Solvers

预生成多难度的PDE数据以用于少样本神经PDE求解器

Naman Choudhary, Vedant Singh, Ameet Talwalkar, Nicholas Matthew Boffi, Mikhail Khodak, Tanya Marwah

机构 * Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) Department of Computer Sciences, UW-Madison(威斯康星大学麦迪逊分校计算机科学系)

专题命中 预训练与数据 :foundation model(abstract);分类 cs.LG

AI总结 通过预生成多难度的PDE数据,提升少样本神经PDE求解器的性能与效率。

Comments 10 Pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01686 2026-01-26 physics.comp-ph 50%

A Graph Neural Network for the Era of Large Atomistic Models

面向大规模原子模型时代的图神经网络

Duo Zhang, Anyang Peng, Chun Cai, Wentao Li, Yuanchang Zhou, Jinzhe Zeng, Mingyu Guo, Chengqian Zhang, Bowen Li, Hong Jiang, Tong Zhu, Weile Jia, Linfeng Zhang, Han Wang

专题命中 预训练与数据 :foundation model(abstract)

AI总结 DPA3是一种基于线图序列的图神经网络,针对大规模原子模型时代设计,通过参数扩展和数据集编码机制,在多个基准任务中表现出卓越的泛化能力和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17406 2026-01-26 cs.CR cs.IR 50%

ProveRAG: Provenance-Driven Vulnerability Analysis with Automated Retrieval-Augmented LLMs

ProveRAG: 基于溯源的漏洞分析与自动化检索增强大语言模型

Reza Fayyazi, Stella Hoyos Trueba, Michael Zuzak, Shanchieh Jay Yang

专题命中 预训练与数据 :LLM(abstract)

AI总结 ProveRAG通过自动化检索增强和自我批评机制,提升大语言模型在漏洞分析中的准确性和时效性,提供高精度的漏洞利用和缓解策略。

Journal ref IEEE Access, vol. 13, pp. 212815-212826, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏