arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12486 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12486 篇

2412.11084 2024-12-17 cs.LG q-bio.GN q-bio.QM 72%

BarcodeMamba: State Space Models for Biodiversity Analysis

Tiancheng Gao, Graham W. Taylor

专题命中 预训练与数据 :foundation model(abstract,comments);pretraining(abstract);分类 cs.LG

Comments 9 pages, 2 figures, accepted at Foundation Models for Science: Progress, Opportunities, and Challenges Workshop (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07420 2023-12-13 cs.LG cs.CY 72%

FairSISA: Ensemble Post-Processing to Improve Fairness of Unlearning in LLMs

Swanand Ravindra Kadhe, Anisa Halimi, Ambrish Rawat, Nathalie Baracaldo

专题命中 预训练与数据 :language model(abstract,comments);large language model(abstract);分类 cs.LG

Comments Accepted in NeurIPS 2023 Workshop on Socially Responsible Language Modelling Research (SoLaR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26917 2026-08-25 cs.CV 版本更新 71%

AnimateAnyMesh++: A Flexible Feed-Forward Framework for High-Fidelity Text-Driven Mesh Animation

AnimateAnyMesh++: 一种灵活的4D基础模型用于高质量文本驱动的网格动画

Zijie Wu, Chaohui Yu, Fan Wang, Xiang Bai

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) DAMO Academy, Alibaba Group(阿里巴巴达摩院) Hupan Lab, Hangzhou, China(湖畔实验室) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)

专题命中 预训练与数据 :foundation model(title)

AI总结 本文提出AnimateAnyMesh++,通过扩展数据集、改进架构和生成能力,实现高质量文本驱动的网格动画,提升了轨迹重建和几何保真度。

Comments 15 pages, TPAMI 2026 accepted, code url: https://github.com/JarrentWu1031/AnimateAnyMesh-pp

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03500 2026-08-05 cs.CY 新提交 71%

LLM-Assisted Review Prioritization for German Statutory Health Insurance Websites: A Multi-Stage Corpus Audit

大语言模型辅助的德国法定医疗保险网站审核优先级排序:多阶段语料审计

Martin Möller

专题命中 预训练与数据 :LLM(title)

AI总结 本研究针对德国法定医疗保险网站,提出一种结合多步骤的大语言模型辅助审核优先级排序工作流程,经实验验证可有效分配审核工作量,区分AI来源信号与质量声明。

Comments 31 pages, 5 figures. Also archived at Zenodo: https://doi.org/10.5281/zenodo.21738595

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02707 2026-07-09 cs.CV 新提交 71%

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

VLRC:视觉语言重投影一致性作为用于更好的前馈3D预训练的可扩展信号

Marwane Hariat, David Filliat, Antoine Manzanera

机构 * U2IS, ENSTA – Institut Polytechnique de Paris(U2IS,法国国立高等先进技术学校 - 巴黎综合理工学院) Pôle Recherche, Agence Ministérielle pour l’IA de Défense(军事人工智能部研究中心)

专题命中 预训练与数据 :pretraining(title)

AI总结 研究提出视觉语言重投影一致性(VLRC),利用冻结的视觉语言表征作为语义多视图监督,无需额外3D标注,能提升3D重建精度等,为前馈3D预训练提供可扩展辅助目标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05906 2026-07-08 cs.CV 新提交 71%

GaussFusion: Towards Multimodal 3D Gaussian Pretraining

高斯融合:迈向多模态3D高斯预训练

Zhixuan You, Jihua Zhu, Yiding Sun, Zihao Guo, Haozhe Cheng, Dongxu Zhang, Lin Chen, Hainan Luo

机构 * Wuhu HIT Robot Technology Research Institute Co., Ltd.(芜湖哈工大机器人技术研究院有限公司)

专题命中 预训练与数据 :pretraining(title)

AI总结 研究提出GaussFusion多模态3D高斯预训练框架,通过跨模态语义对齐集成图像和文本监督,还提出高斯显著性引导的多尺度空洞掩码,实验表明该方法提高了高斯表示可迁移性,在ModelNet40和ScanObjectNN上有性能提升。

Comments 32 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25546 2026-06-25 cs.CV 新提交 71%

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

面向3D计算机断层扫描的疾病中心视觉语言预训练与混合视觉编码

Bowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang, Wenrui Dai, Hongkai Xiong, Ling Zhang, Jianpeng Zhang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab, 310023, Hangzhou, China(虎扑实验室,杭州,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Zhejiang University, China(浙江大学,中国) Fudan University, China(复旦大学,中国)

专题命中 预训练与数据 :pretraining(title)

AI总结 提出一种结合CNN-ViT混合编码器、疾病级对比学习和诊断感知提示的视觉语言预训练框架,在CT-RATE和Rad-ChestCT上取得最优性能,并提升零样本诊断可靠性。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04898 2026-06-04 cs.CV 71%

CDPM-Align: Multi-Scale Guidance-Aligned Diffusion Pretraining for Robust Few-Shot Anatomical Landmark Detection

CDPM-Align:用于鲁棒少样本解剖标志检测的多尺度引导对齐扩散预训练

Roberto Di Via, Irina Voiculescu, Francesca Odone, Vito Paolo Pastore

机构 * MaLGa DIBRIS, University of Genoa(DIBRIS,热那亚大学) University of Genoa(热那亚大学) Department of Computer Science, University of Oxford(奥大利大学计算机科学系)

专题命中 预训练与数据 :pretraining(title)

AI总结 提出多尺度引导对齐的条件扩散预训练方法CDPM-Align,通过生成式预训练学习鲁棒表示,在少样本和低标注场景下提升解剖标志检测的准确性和不确定性。

Comments Accepted MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23840 2026-05-25 cs.CV 71%

MuellerPT: Decomposition Driven Pretraining for Dense Learning in Mueller Polarimetry

MuellerPT: 穆勒偏振测量中密集学习的分解驱动预训练

Adam Tlemsani, Yingdian Li, Maxime Giot, Naim Slim, Christopher J. Peters, Abhijeet Ghosh, Daniel S. Elson

机构 * Department of Computing, Imperial College London(帝国理工学院计算机系) Hamlyn Centre for Robotic Surgery, Imperial College London(帝国理工学院机器人外科中心) Department of Surgery and Cancer, Imperial College London(帝国理工学院外科与癌症系) Xi’an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences(中国科学院西安光学精密机械研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 预训练与数据 :pretraining(title)

AI总结 提出MuellerPT,一种通过预测Lu-Chipman分解图进行物理引导预训练的方法,在少样本分割和分类任务中显著提升标签效率和跨样本泛化能力。

Comments Accepted to 29th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19202 2026-05-08 cs.CV 71%

UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models

UniE2F: 一种基于视频基础模型的统一扩散框架用于事件到帧重建

Gang Xu, Zhiyu Zhu, Junhui Hou

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) Department of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学系)

专题命中 预训练与数据 :foundation model(title)

AI总结 本文提出UniE2F框架,利用预训练视频扩散模型的生成先验,从稀疏事件数据中重建高保真视频帧,通过引入事件基帧间残差引导提升重建精度,并扩展到零样本视频帧插值与预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01846 2026-04-30 cs.CV 71%

Investigating the Segment Anything Foundation Model for Mapping Smallholder Agriculture Field Boundaries Without Training Labels

探究无需标注的Segment Anything基础模型用于映射小农户农业田块边界

Pratyush Tripathy, Kathy Baylis, Kyle Wu, Jyles Watson, Ruizhe Jiang

机构 * Department of Geography, University of California, Santa Barbara, CA, 93106 USA(地理系,加州大学圣巴巴拉分校,加州,圣巴巴拉,93106 USA)

专题命中 预训练与数据 :foundation model(title)

AI总结 本文探讨了使用无需额外训练的Segment Anything模型(SAM)在印度比哈尔州利用SkySat影像映射小农户农业田块边界,评估了不同模型版本、输入尺寸和多日期影像对精度的影响,结果显示SAM在无标注数据下能准确识别约58%的田块边界。

Comments 11 pages, 6 main figures, 7 supplementary figures

Journal ref Science of Remote Sensing, Volume 13, 2026, 100425

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03222 2026-03-16 cs.CV 71%

Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining

Mocap-2-to-3:多视角提升用于单目运动恢复的2D预训练

Zhumei Wang, Zechen Hu, Ruoxi Guo, Huaijin Pi, Ziyong Feng, Liang Zhang, Mingtao Pei, Siyuan Huang

机构 * Beijing Institute of Technology(北京理工大学) State Key Laboratory of General Artificial Intelligence, BIGAI(国家一般人工智能重点实验室,BIGAI) Deep Glint Zhejiang University(浙江大学) The University of Hong Kong(香港大学) Shandong Agricultural University(山东农业大学)

专题命中 预训练与数据 :pretraining(title)

AI总结 本文提出Mocap-2-to-3框架,通过多视角合成过程和分阶段训练提升单目运动恢复的精度与泛化能力,实现更精确的物理空间定位。

Comments Project page: https://wangzhumei.github.io/mocap-2-to-3/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09627 2026-03-11 eess.AS 71%

Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models

Speech-Omni-Lite: 用于视觉-语言模型的便携式语音接口

Dehua Tao, Xuan Luo, Daxin Tan, Kai Chen, Lanqing Hong, Jing Li, Ruifeng Xu, Xiao Chen

专题命中 预训练与数据 :language model(title)

AI总结 Speech-Omni-Lite通过轻量模块扩展视觉-语言模型,实现低成本语音理解和生成,同时保持原有性能,实验显示其在少量语音数据下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00732 2026-03-03 cs.RO cs.CV 71%

UniHM: Unified Dexterous Hand Manipulation with Vision Language Model

UniHM: 一体化的视觉语言模型用于统一的灵巧手操作

Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) InstAdapt

专题命中 预训练与数据 :language model(title)

AI总结 UniHM通过统一的视觉语言模型实现灵巧手操作,利用开放词汇指令提升泛化能力和物理可行性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09333 2026-02-06 cs.CV cond-mat.mtrl-sci eess.IV 71%

MaskTerial: A Foundation Model for Automated 2D Material Flake Detection

MaskTerial: 一种用于自动2D材料片状物检测的基础模型

Jan-Lucas Uslu, Alexey Nekrasov, Alexander Hermans, Bernd Beschoten, Bastian Leibe, Lutz Waldecker, Christoph Stampfer

机构 * nd Institute of Physics(第二物理研究所) JARA-FIT, RWTH Aachen University, 52074 Aachen, Germany(JARA-FIT,亚琛工业大学) Visual Computing Institute, RWTH Aachen University, 52074 Aachen, Germany(视觉计算研究所,亚琛工业大学)

专题命中 预训练与数据 :foundation model(title)

AI总结 MaskTerial是一种利用实例分割网络和不确定性估计模型,实现自动2D材料片状物检测的基础模型,显著提升了低对比度材料的检测性能。

Comments 9 pages, 5 figures

Journal ref Digital Discovery 4, 3744 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05385 2026-01-12 cs.SE 71%

DafnyPro: LLM-Assisted Automated Verification for Dafny Programs

DafnyPro:基于LLM的自动验证框架用于Dafny程序

Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche

专题命中 预训练与数据 :LLM(title)

AI总结 DafnyPro通过增强LLM生成验证注释的能力,提升了Dafny程序的自动验证效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15967 2025-11-21 cs.CV 71%

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

InfoCLIP: 通过信息论对齐转移连接视觉语言预训练与开放词汇语义分割

Muyao Yuan, Yuanhong Zhang, Weizhan Zhang, Lan Ma, Yuan Gao, Jiangyong Ying, Yudeng Xin

专题命中 预训练与数据 :pretraining(title)

AI总结 InfoCLIP通过信息论对齐转移提升开放词汇语义分割的性能,有效解决预训练CLIP在微调过程中的过拟合问题。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14270 2025-11-11 cs.CV cs.GR 71%

GauSSmart: Enhanced 3D Reconstruction through 2D Foundation Models and Geometric Filtering

Alexander Valverde, Brian Xu, Yuyin Zhou, Meng Xu, Hongyun Wang

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) Brown University(布朗大学) Kean University(凯恩大学)

专题命中 预训练与数据 :foundation model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04426 2025-11-07 cs.CV 71%

HideAndSeg: an AI-based tool with automated prompting for octopus segmentation in natural habitats

Alan de Aguiar, Michaella Pereira Andrade, Charles Morphy D. Santos, João Paulo Gois

机构 * Universidade Federal do ABC (UFABC)(巴西圣安德烈大学联邦ABC分校)

专题命中 预训练与数据 :prompting(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08054 2025-10-13 cs.CV 71%

RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models

Moon Ye-Bin, Roy Miles, Tae-Hyun Oh, Ismail Elezi, Jiankang Deng

机构 * POSTECH Huawei London Research Center(华为伦敦研究中心) KAIST(韩国科学技术院) Imperial College London(伦敦帝国理工学院)

专题命中 预训练与数据 :language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08586 2025-09-11 eess.IV cs.CV 71%

CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining

Prashant Singh Basnet, Roshan Chitrakar

机构 * The British College, Keele University(凯利大学英国学院) Nepal College of Information Technology, Pokhara University(尼泊尔信息技术学院,波克拉大学)

专题命中 预训练与数据 :pretraining(title)

Comments 8 pages, 5 Tables, 5 Figures. Manuscript submitted to ICOIICS 2025 Conference. Currently, under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01550 2025-09-04 cs.SE 71%

RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale

Zhilong Chen, Chengzong Zhao, Boyuan Chen, Dayi Lin, Yihao Chen, Arthur Leung, Gopi Krishnan Rajbahadur, Gustavo A. Oliva, Haoxiang Zhang, Aaditya Bhatia, Chong Chun Yong, Ahmed E. Hassan

专题命中 预训练与数据 :SFT(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21960 2025-07-30 cs.CV 71%

PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction

Jiahui Ren, Mochu Xiang, Jiajun Zhu, Yuchao Dai

机构 * School of Electronics and Information, Northwestern Polytechnical University and Shaanxi Key Laboratory of Information Acquisition and Processing(电子工程学院、西北工业大学和陕西省信息获取与处理重点实验室)

专题命中 预训练与数据 :pretraining(title)

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03474 2025-07-01 cs.CV 71%

Multi-encoder nnU-Net outperforms transformer models with self-supervised pretraining

Seyedeh Sahar Taheri Otaghsara, Reza Rahmanzadeh

专题命中 预训练与数据 :pretraining(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17837 2025-06-24 cs.CV 71%

Time-Contrastive Pretraining for In-Context Image and Video Segmentation

Assefa Wahd, Jacob Jaremko, Abhilash Hareendranathan

机构 * Department of Radiology(放射科部门)

专题命中 预训练与数据 :pretraining(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11241 2025-06-19 cs.CV 71%

CooPre: Cooperative Pretraining for V2X Cooperative Perception

Seth Z. Zhao, Hao Xiang, Chenfeng Xu, Xin Xia, Bolei Zhou, Jiaqi Ma

机构 * University of California, Los Angeles(加州大学洛杉矶分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 预训练与数据 :pretraining(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13455 2025-06-17 eess.AS cs.SD 71%

Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling

Wenmiao Gao, Yang Xiao

专题命中 预训练与数据 :pretraining(title)

Comments Technical report for DCASE 2025 Challenge Task 3

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18052 2025-06-04 cs.CV 71%

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

Yue Li, Qi Ma, Runyi Yang, Huapeng Li, Mengjiao Ma, Bin Ren, Nikola Popovic, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Martin R. Oswald, Danda Pani Paudel

机构 * University of Amsterdam(阿姆斯特丹大学) Computer Vision Lab, ETH Zurich(苏黎世联邦理工学院计算机视觉实验室) INSAIT, Sofia University ”St. Kliment Ohridski”(索菲亚大学”圣克莱门特·欧赫里迪斯”研究所) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Pisa(比萨大学) University of Trento(特伦特大学)

专题命中 预训练与数据 :pretraining(title)

Comments Our code, model, and dataset will be released at https://unique1i.github.io/SceneSplat_webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24282 2025-06-02 cs.CV 71%

LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization

Zirui Shang, Xinxiao Wu, Shuo Yang

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology(北京智能信息科技重点实验室,计算机科学与技术学院) Beijing Institute of Technology(北京理工大学) Guangdong Laboratory of Machine Perception and Intelligent Computing(广东机器感知与智能计算实验室) Shenzhen MSU-BIT University(深圳MSU-BIT大学)

专题命中 预训练与数据 :LLM(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10847 2025-05-23 cs.CV 71%

Enhancement-Driven Pretraining for Robust Fingerprint Representation Learning

Ekta Gavas, Kaustubh Olpadkar, Anoop Namboodiri

机构 * Centre for Visual Information Technology, International Institute of Information Technology, Hyderabad, India(视觉信息技术中心,国际信息学院,印度海得拉巴) Stony Brook University, USA(石溪大学)

专题命中 预训练与数据 :pretraining(title)

Comments 8 pages, 4 figures, Accepted at 19th VISIGRAPP 2024: VISAPP conference

Journal ref Proceedings of the 19th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP 2024) - Volume 2: VISAPP, ISBN 978-989-758-679-8, ISSN 2184-4321, pages 821-828

详情

展开后加载摘要…

URL PDF HTML 收藏