arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7608 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7608 篇

2511.21580 2025-11-27 cs.SD 50%

Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension

基于谐波- percussive 分离的神经音频编码器带宽扩展方法

Benoît Giniès, Xiaoyu Bie, Olivier Fercoq, Gaël Richard

机构 * LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI、巴黎电信学院、巴黎理工学院)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出一种基于谐波- percussive 分离的神经音频编码器,通过将带宽扩展视为音频令牌预测问题,提升音频信号的高频重构质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20366 2025-11-27 cs.CV 50%

VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the Wild

VGGTFace: 在真实世界中实现拓扑一致的面部几何重建

Xin Ming, Yuxuan Han, Tianyu Huang, Feng Xu

机构 * Xin Ming, Yuxuan Han, Tianyu Huang, Feng Xu(作者)

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 VGGTFace通过引入VGGT和Pixel3DMM实现高效拓扑一致的面部几何重建,适用于真实世界多视角图像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20167 2025-11-26 cs.MM 50%

FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation

FINE: 通过互信息估计进行因子化多模态情感分析

Yadong Liu, Shangfei Wang

专题命中 知识编辑与模型理解 :prompting(abstract)

AI总结 本文提出了一种基于互信息估计的因子化多模态情感分析框架,通过分解模态为共享和独特表示,抑制噪声并提升情感表示质量,从而在多个数据集上优于现有方法。

Comments 15 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17964 2025-11-26 cs.CV 50%

X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification

X-ReID:基于视频的可见-红外人重识别中的多粒度信息交互

Chenyang Yu, Xuehu Liu, Pingping Zhang, Huchuan Lu

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 X-ReID通过多粒度信息交互和跨模态特征学习,提升视频中可见-红外人重识别的性能。

Comments Accepted by AAAI2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19315 2025-11-25 cs.RO 50%

Rethinking Intermediate Representation for VLM-based Robot Manipulation

重新思考基于VLM的机器人操作中的中间表示

Weiliang Tang, Jialin Gao, Jia-Hui Pan, Gang Wang, Li Erran Li, Yunhui Liu, Mingyu Ding, Pheng-Ann Heng, Chi-Wing Fu

机构 * CUHK(香港中文大学) Amazon(亚马逊) UNC(北卡罗来纳大学教堂山分校)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出SEAM表示方法,通过分解中间表示为词汇和语法,提升VLM在机器人操作中的可理解和通用性,结合检索增强的少样本学习策略实现高效操作,并在动作通用性和VLM可理解性上展示出优于主流方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17904 2025-11-25 cs.CV cs.RO 50%

CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation

CUS-GS: 一种紧凑的统一结构高斯点扩散框架用于多模态场景表示

Yuhang Ming, Chenxin Fang, Xingyuan Yu, Fan Zhang, Weichen Dai, Wanzeng Kong, Guofeng Zhang

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院) CAD & CG, Zhejiang University(浙江大学计算机辅助设计与图形学研究所) School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 CUS-GS通过统一结构化高斯点扩散框架,实现多模态场景表示的高效建模与语义一致性。

Comments 15 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17462 2025-11-24 q-fin.PM 50%

Scaling Conditional Autoencoders for Portfolio Optimization via Uncertainty-Aware Factor Selection

通过不确定性感知因子选择实现条件自编码器的规模扩展以用于投资组合优化

Ryan Engel, Yu Chen, Pawel Polak, Ioana Boier

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 本文提出通过不确定性感知因子选择扩展条件自编码器,提升投资组合优化的风险调整后绩效。

Comments 9 pages, 6 figures. Published in Proceedings of the 6th ACM International Conference on AI in Finance (ICAIF '25)

Journal ref ICAIF '25: Proceedings of the 6th ACM International Conference on AI in Finance, pages 123-131, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06958 2025-11-20 cs.CV 50%

Learning from the Right Patches: A Two-Stage Wavelet-Driven Masked Autoencoder for Histopathology Representation Learning

Raneen Younis, Louay Hamdi, Lukas Chavez, Zahra Ahmadi

机构 * PLRI Medical Informatics Institute(PLRI医学信息学研究所) CAIMed Reserch Center(CAIMed研究中心) Hannover Medical School(汉诺威医学院) Computer Science Institute(计算机科学研究所) Leibniz University Hannover(莱比锡大学汉诺威分校) Sanford Burnham Prebys Medical Discovery Institute(桑福医疗发现研究所) Rady Children’s Institute for Genomic Medicine(拉迪儿童基因医学研究所) University of California San Diego(加州大学圣地亚哥分校)

专题命中 知识编辑与模型理解 :pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13586 2025-11-20 cs.CV 50%

Adaptive Multi-Scale Integration Unlocks Robust Cell Annotation in Histopathology Images

Yinuo Xu, Yan Cui, Mingyao Li, Zhi Huang

机构 * Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学) Department of Bioengineering, University of Pennsylvania(生物工程系,宾夕法尼亚大学) Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学) Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)

专题命中 知识编辑与模型理解 :foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13881 2025-11-19 cs.CV 50%

VLMs Guided Interpretable Decision Making for Autonomous Driving

Xin Hu, Taotao Jing, Renran Tian, Zhengming Ding

机构 * Department of Computer Science, Tulane University(路易斯安那大学计算机科学系) Qualcomm(高通公司) Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系)

专题命中 知识编辑与模型理解 :language model(abstract)

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20105 2025-11-19 cs.CV 50%

FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models

Barbara Toniella Corradini, Mustafa Shukor, Paul Couairon, Guillaume Couairon, Franco Scarselli, Matthieu Cord

机构 * DIISM University of Siena(DIISM锡耶纳大学) CNRS, ISIR Sorbonne University(CNRS,ISIR索邦大学) Inria, ARCHES Sorbonne University(Inria,ARCHES索邦大学) Valeo.ai Sorbonne University(Valeo.ai索邦大学)

专题命中 知识编辑与模型理解 :foundation model(abstract)

Journal ref Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12978 2025-11-18 cs.CV 50%

Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach

Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi

机构 * CVIT, Kohli Centre for Intelligent Systems, IIIT Hyderabad(IIIT海得拉尔计算机视觉研究所、Kohli智能系统中心) Adobe Research, Bengaluru(Adobe研究)

专题命中 知识编辑与模型理解 :language model(abstract)

Comments 25 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18934 2025-11-17 cs.CR 50%

Revealing Adversarial Smart Contracts through Semantic Interpretation and Uncertainty Estimation

Yating Liu, Xing Su, Hao Wu, Sijin Li, Yuxi Cheng, Fengyuan Xu, Sheng Zhong

专题命中 知识编辑与模型理解 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05680 2025-11-11 cs.RO 50%

VLM-driven Skill Selection for Robotic Assembly Tasks

Jeong-Jung Kim, Doo-Yeol Koh, Chang-Hyun Kim

机构 * Department of AI Machinery, Korea Institute of Machinery & Materials(人工智能机械系,韩国机械材料研究院)

专题命中 知识编辑与模型理解 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05432 2025-11-10 cs.CV 50%

Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis

Dogucan Yaman, Seymanur Akti, Fevziye Irem Eyiokur, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11350 2025-11-10 cs.RO 50%

Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild

Derek Ming Siang Tan, Shailesh, Boyang Liu, Alok Raj, Qi Xuan Ang, Weiheng Dai, Tanishq Duhan, Jimmy Chiun, Yuhong Cao, Florian Shkurti, Guillaume Sartoretti

专题命中 知识编辑与模型理解 :language model(abstract)

Comments Accepted for presentation at CORL 2025. Code, models, and data are available at https://search-tta.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04078 2025-11-07 cs.CV 50%

Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment

Zehui Feng, Chenqi Zhang, Mingru Wang, Minuo Wei, Shiwei Cheng, Cuntai Guan, Ting Han

专题命中 知识编辑与模型理解 :pretraining(abstract)

Comments 30 pages, 16 figures, under review as a conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26799 2025-10-31 cs.CV 50%

Masked Diffusion Captioning for Visual Feature Learning

Chao Feng, Zihao Wei, Andrew Owens

机构 * University of Michigan(密歇根大学) Cornell University(康奈尔大学) University of Maryland(马里兰大学)

专题命中 知识编辑与模型理解 :language model(abstract)

Comments EMNLP 2025 (Findings). Project page: https://cfeng16.github.io/mdlm4vfl/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22322 2025-10-28 cs.CV 50%

Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning

Ali Javidani, Babak Nadjar Araabi, Mohammad Amin Sadeghi

机构 * School of Electrical and Computer Engineering, College of Engineering, University of Tehran(埃塞电气与计算机工程学院,工程学院,德黑兰大学) Qatar Computing Research Institute, Hamad bin Khalifa University(卡塔尔计算研究所,哈马德·本·卡伊夫大学)

专题命中 知识编辑与模型理解 :pretraining(abstract)

Comments Accepted in IEEE Signal Processing Letters, 2025

Journal ref IEEE Signal Processing Letters, vol. 32, pp. 3730-3734, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10097 2025-10-28 cs.CV 50%

Gesplat: Robust Pose-Free 3D Reconstruction via Geometry-Guided Gaussian Splatting

Jiahui Lu, Haihong Xiao, Xueyan Zhao, Wenxiong Kang

专题命中 知识编辑与模型理解 :foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18703 2025-10-22 cs.CV 50%

Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents

Yiqi Lin, Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学展示实验室) Central South University(中南大学) Microsoft(微软公司)

专题命中 知识编辑与模型理解 :language model(abstract)

Comments Project page: this https://linyq17.github.io/VC2L/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18321 2025-10-22 cs.CV 50%

Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding

Jinlin Li, Yuran Wang, Yifei Yuan, Xiao Zhou, Yingying Zhang, Xixian Yong, Yefeng Zheng, Xian Wu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallup 学院) Department of Electrical and Computer Engineering, McGill University(麦吉尔大学电气与计算机工程系) School of Statistics, Renmin University of China(中国人民大学统计学院) Tencent Jarvis Lab(腾讯 Jarvis 实验室) Medical Artificial Intelligence Lab, Westlake University(西湖大学医学人工智能实验室)

专题命中 知识编辑与模型理解 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14671 2025-10-21 cs.CV 50%

UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens

Ruichuan An, Sihan Yang, Renrui Zhang, Zijun Shen, Ming Lu, Gaole Dai, Hao Liang, Ziyu Guo, Shilin Yan, Yulin Luo, Bocheng Zou, Chaoqun Yang, Wentao Zhang

机构 * Peking University(北京大学) Xi’an JiaoTong University(西安交通大学) CUHK(香港中文大学) Intel Labs, China(中国英特尔实验室) Nanjing University(南京大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Tsinghua University(清华大学)

专题命中 知识编辑与模型理解 :language model(abstract)

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13809 2025-10-16 cs.CV 50%

PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning

Sihui Ji, Xi Chen, Xin Tao, Pengfei Wan, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 知识编辑与模型理解 :preference optimization(abstract)

Comments Project Page: https://sihuiji.github.io/PhysMaster-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12400 2025-10-16 cs.IR 50%

QUIDS: Query Intent Description for Exploratory Search via Dual Space Modeling

Yumeng Wang, Xiuying Chen, Suzan Verberne

专题命中 知识编辑与模型理解 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12395 2025-10-15 cs.CR 50%

IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion

Ye Tian, Yanqiu Yu, Liangliang Song, Zhiquan Liu, Yanbin Wang, Jianguo Sun

专题命中 知识编辑与模型理解 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10524 2025-10-14 cs.CV 50%

Unified Open-World Segmentation with Multi-Modal Prompts

Yang Liu, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen, Yuling Xi, Bo Feng, Hao Wang, Shiyu Li, Chunhua Shen

机构 * Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University of Technology(浙江工业大学) Apple(苹果公司)

专题命中 知识编辑与模型理解 :foundation model(abstract)

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10466 2025-10-14 cs.CV 50%

When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance

Jinjin Cao, Zhiyang Chen, Zijun Wang, Liyuan Ma, Weijian Luo, Guojun Qi

机构 * MAPLE Lab, Westlake University(西溪大学MAPLE实验室)

专题命中 知识编辑与模型理解 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07839 2025-10-10 cs.CV 50%

AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views

Yijie Gao, Houqiang Zhong, Tianchi Zhu, Zhengxue Cheng, Qiang Hu, Li Song

机构 * School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China(信息科学与电子工程学院,上海交通大学,上海,中国) Cooperative Mediant Innovation Center, Shanghai Jiao Tong University, Shanghai, China(协同医疗创新中心,上海交通大学,上海,中国) SJTU Paris Elite Institute of Technology, Shanghai Jiao Tong University, Shanghai, China(上海交通大学巴黎精英技术学院)

专题命中 知识编辑与模型理解 :foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07417 2025-10-10 cs.RO 50%

FLEET: Formal Language-Grounded Scheduling for Heterogeneous Robot Teams

Corban Rivera, Grayson Byrd, Meghan Booker, Bethany Kemp, Allison Gaines, Emma Holmes, James Uplinger, Celso M de Melo, David Handelman

机构 * JHU APL(约翰霍普金斯大学应用物理实验室) JHU(约翰霍普金斯大学) DEVCOM ARL(国防高级研究计划局)

专题命中 知识编辑与模型理解 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏