arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-15 至 2025-12-15 共收录 8 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 8 篇

2502.05790 2025-12-15 cs.LG 90%

Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining

打破冻结子空间:用于大语言模型预训练低秩优化的重要性采样

Haochen Zhang, Junze Yin, Guanchu Wang, Zirui Liu, Lin F. Yang, Tianyi Zhang, Anshumali Shrivastava, Vladimir Braverman

机构 * Rice University(Rice大学) University of North Carolina at Charlotte(北卡罗来纳州立大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) University of California, Los Angeles(加州大学洛杉矶分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 预训练与数据 :LLM(title,abstract);pretraining(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出了一种具有可证明收敛性的低秩优化方法,用于大语言模型预训练,以解决主导子空间在预训练中冻结的问题,并在实验中表现出优越的性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05578 2025-12-15 cs.LG cs.CL cs.CR 85%

The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation

LLM记忆景观:机制、测量与缓解

Alexander Xiong, Xuandong Zhao, Aneesh Pappu, Dawn Song

机构 * UC Berkeley(伯克利大学) Google DeepMind(谷歌DeepMind)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文研究了LLM记忆现象的机制、测量方法及缓解策略,探讨了影响记忆的因素和检测技术,并分析了其法律伦理影响及缓解措施。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17186 2025-12-15 cs.CV 78%

Advancing Weakly-Supervised Change Detection in Satellite Images via Adversarial Class Prompting

通过对抗性类别提示推进卫星图像弱监督变化检测

Zhenghui Zhao, Chen Wu, Di Wang, Hongruixuan Chen, Cuiqun Chen, Zhuo Zheng, Bo Du, Liangpei Zhang

机构 * State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(信息工程测绘遥感国家重点实验室,武汉大学) School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学) Graduate School of Frontier Sciences, University of Tokyo(前沿科学研究院,东京大学) Institute of Geodesy and Photogrammetry, ETH Zürich(测绘学研究院,苏黎世联邦理工学院) Department of Computer Science, Stanford University(计算机科学系,斯坦福大学)

专题命中 预训练与数据 :prompting(title,abstract)

AI总结 提出对抗性类别提示方法,通过对抗性扰动和原型校正提升卫星图像弱监督变化检测性能。

Comments Accepted by IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11458 2025-12-15 cs.CV cs.AI 77%

Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation

通过无训练测试时适应提升基于骨架的零样本动作识别

Jingmin Zhu, Anqi Zhu, Hossein Rahmani, Jun Liu, Mohammed Bennamoun, Qiuhong Ke

机构 * Monash University(墨尔本大学) Lancaster University(兰卡斯特大学) University of Western Australia(西澳大学)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 通过引入Skeleton-Cache框架,利用LLM引导的语义先验实现无训练测试时适应,提升基于骨架的零样本动作识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05103 2025-12-15 cs.LG cs.AI cs.CL cs.CV 67%

TV2TV: A Unified Framework for Interleaved Language and Video Generation

TV2TV:一种用于交错语言和视频生成的统一框架

Xiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen, Karthik Padthe, Liam Robbins, Amir Bar, Delong Chen, Michal Drozdzal, Maha Elbayad, Yushi Hu, Shang-Wen Li, Sreya Dutta Roy, Jakob Verbeek, XuDong Wang, Marjan Ghazvininejad, Luke Zettlemoyer, Emily Dinan

机构 * Meta FAIR

专题命中 预训练与数据 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 TV2TV通过统一框架实现语言与视频生成的交错过程,提升视频生成的视觉质量和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11567 2025-12-15 cs.CL cs.MM 57%

Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet

扩展议会语料库以包含议员的推文:利用MultiParTweet进行自动标注和评估

Mevlüt Bagci, Ali Abusaleh, Daniel Baumartz, Giueseppe Abrami, Maxim Konca, Alexander Mehler

专题命中 预训练与数据 :language model(abstract);分类 cs.CL

AI总结 本文提出MultiParTweet,通过自动标注和人工验证,整合了多语言推文及媒体内容,展示了模型间互为预测性及多模态注释的优越性。

Comments Submitted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02146 2025-12-15 cs.CL cs.SD eess.AS 57%

Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation

可唱旋律到歌词生成中的词汇与格式联合学习

Longshen Ou, Xichu Ma, Ye Wang

专题命中 预训练与数据 :pretraining(abstract);分类 cs.CL

AI总结 本文提出通过联合学习词汇和格式来提升旋律到歌词生成的可唱性,实验显示在行数和音节数要求上表现更优。

Comments An extension of our previous work arXiv:2305.16816 [cs.CL]

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11596 2025-12-15 physics.chem-ph 50%

Transfer learning of GW-Bethe-Salpeter Equation excitation energies

GW-贝瑟-萨尔皮特方程激发能的迁移学习

Dario Baum, Arno Förster, Lucas Visscher

专题命中 预训练与数据 :pretraining(abstract)

AI总结 本文通过迁移学习方法,利用预训练的图神经网络在有限的高保真数据下提升对准粒子和激发能的预测精度,扩展多体预测的化学空间范围。

详情

展开后加载摘要…

URL PDF HTML 收藏