Revisiting Theory of Contrastive Learning for Domain Generalization
重新审视对比学习在领域泛化中的理论
Ali Alvandi, Mina Rezaei
机构
*
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Department of Computer Science, Sharif University of Technology(技术学院计算机科学系)
;
Department of Statistics, LMU Munich(慕尼黑大学统计系)
Comments22 pages, 4 figures. Related to our earlier preprint "The brain versus AI" (arXiv:2411.16075) but a distinct article. The earlier work surveyed broad brain-AI parallels; here we focus on world-model-based computation and convergent evolution between the brain and AI, especially large language models
From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models
从词向量到多模态嵌入:大型语言模型的技术、应用与未来方向
Charles Zhang, Benji Peng, Xintian Sun, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Yichao Zhang, Xinyuan Song, Cheng Fei, Caitlyn Heqi Yin, Lawrence KQ Yan, Hongyang He, Tianyang Wang
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
Simon Fraser University(西蒙弗雷泽大学)
;
Kyoto University(京都大学)
;
National Taiwan Normal University(台湾师范大学)
;
Purdue University(普渡大学)
;
The University of Texas at Dallas(德克萨斯大学达拉斯分校)
;
Emory University(埃默里大学)
;
Cornell University(康奈尔大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
The Hong Kong University of Science(香港科学大学)
;
University of Liverpool(利物浦大学)
;
University of Warwick(沃里克大学)
专题命中
知识编辑与模型理解
:language model(title,abstract);large language model(title);分类 cs.CL
Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
几何不确定性用于检测和纠正大语言模型中的幻觉
Edward Phillips, Sean Wu, Soheila Molaei, Danielle Belgrave, Anshul Thakur, David Clifton
机构
*
Department of Engineering Science, University of Oxford(牛津大学工程科学系)
;
GlaxoSmithKline(葛兰素史克)
;
Oxford Suzhou Centre for Advanced Research(牛津苏黎世高级研究中心)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL
CommentsRevision. Clarified positioning as a unified geometric framework for global and local uncertainty in LLMs. Added baselines (Degree, Eccentricity) and expanded comparison to related methods. Included ablations (PCA dimension, number of archetypes, number of samples) and complexity analysis. Extended discussion of medical QA results and model-specific behaviour
CommentsThis paper is accepted to the New Ideas and Emerging Results (NIER) track of the ACM/IEEE 28th International Conference on Model Driven Engineering Languages and Systems (MODELS)
CommentsCamera-ready version. Oral presentation at IJCNLP-AACL 2025 (14th International Joint Conference on Natural Language Processing and 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics), Mumbai, India, December 20-24, 2025
机构
*
Computer Engineering K J Somaiya Institute of Technology Mumbai, India
;
Data Science K J Somaiya Institute of Technology Mumbai, India
;
Information Technology K J Somaiya Institute of Technology Mumbai, India