arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12321 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 12321 篇

1307.1662 2014-06-30 cs.CL cs.LG 62%

Polyglot: Distributed Word Representations for Multilingual NLP

Rami Al-Rfou, Bryan Perozzi, Steven Skiena

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG

Comments 10 pages, 2 figures, Proceedings of Conference on Computational Natural Language Learning CoNLL'2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1211.2290 2012-11-13 cs.CL cs.AI 62%

Dating Texts without Explicit Temporal Cues

Abhimanu Kumar, Jason Baldridge, Matthew Lease, Joydeep Ghosh

专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12440 2026-08-18 cs.SE cs.AI 版本更新 61%

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

采用AI编码智能体的规范优先收敛:在71.7万行代码库中拆解189个文件的核心架构不变量的案例研究

Joel Abenhaim

专题命中 其他LLM :language model(abstract);分类 cs.AI;LLM(comments)

AI总结 本文通过案例研究,展示AI编码智能体在规范优先协议下,成功在71.7万行代码库中拆解核心架构不变量,耗时3天、成本2430美元,修正201个缺陷,涉及189个文件。

Comments 14 pages, 4 figures, 3 tables. v2: added plain-text log URLs in Section 10 for LLM readability

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08005 2026-08-11 cs.LG 版本更新 61%

Preference Redirection via Attention Concentration: An Attack on Computer Use Agents

通过注意力集中进行偏好重定向:对计算机使用代理的一种攻击

Dominik Seip, Matthias Hein

机构 * Tübingen AI Center, University of Tübingen(图宾根大学图宾根人工智能中心)

专题命中 其他LLM :foundation model(abstract);分类 cs.LG;language model(journal_ref)

AI总结 本文提出PRAC攻击,通过重定向模型注意力操纵CUA选择过程,揭示了视觉模态的漏洞,威胁到基于开放权重模型的CUA安全。

Journal ref Conference on Language Modeling (COLM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01637 2026-05-05 cs.LG cs.CC cs.DM math.CO 61%

The Banach-Butterfly Invariant: Influence-Adaptive Walsh Geometry for Ternary Polynomial Threshold Functions

Banach-Butterfly 不变量:适应影响的Walsh几何用于三元多项式阈值函数

Gorgi Pavlov

机构 * Lehigh University(莱德大学) Johnson and Johnson(强生公司)

专题命中 其他LLM :LLM(abstract_cn,comments);分类 cs.LG

AI总结 本文提出Banach-Butterfly不变量,用于三元多项式阈值函数的适应影响的Walsh几何分析,通过影响向量的Schur凸性分离函数,并证明其作为收缩不变量的性质。

Comments 21 pages, 3 figures. Theory paper; LLM-application companion in preparation. Code, certificates, and 616,126 NPN-canonical n=5 representatives in supplementary repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11389 2025-11-17 cs.CL 61%

Studies with impossible languages falsify LMs as models of human language

Jeffrey S. Bowers, Jeff Mitchell

专题命中 其他LLM :language model(abstract,comments);分类 cs.CL

Comments Commentary on Futrell, R., & Mahowald, K. arXiv:2501.17047 (in press). How linguistics learned to stop worrying and love the language models. Behavioural and Brain Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01845 2025-10-03 cs.CL cs.CV 61%

Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models

Ece Takmaz, Lisa Bylinina, Jakub Dotlacil

机构 * Utrecht University(乌特勒支大学)

专题命中 其他LLM :language model(abstract,comments);分类 cs.CL

Comments Accepted to the EMNLP 2025 workshop BabyLM: Accelerating language modeling research with cognitively plausible datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21226 2025-07-29 cs.CV cs.AI 61%

MemeBLIP2: A novel lightweight multimodal system to detect harmful memes

Jiaqi Liu, Ran Tong, Aowei Shen, Shuzheng Li, Changlin Yang, Lisha Xu

机构 * Mathematics and Statistics Department, University of Texas at Dallas(德克萨斯大学达拉斯分校数学与统计学系)

专题命中 其他LLM :language model(abstract,comments);分类 cs.AI

Comments 11 pages, 3 figures. Accepted at the First Workshop on Multimodal Knowledge and Language Modeling (MKLM), IJCAI-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01964 2024-06-18 stat.ML cs.LG 61%

Position: Understanding LLMs Requires More Than Statistical Generalization

Patrik Reizinger, Szilvia Ujváry, Anna Mészáros, Anna Kerekes, Wieland Brendel, Ferenc Huszár

专题命中 其他LLM :LLM(abstract,comments);分类 cs.LG

Comments Accepted as a position paper at ICML2024, Code: https://github.com/rpatrik96/llm-non-identifiability

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29313 2026-09-01 cs.CV cs.AI 新提交 57%

Hyper3-CLIP: Hierarchy-Conditioned Hyperbolic Vision-Language Training

Hyper3-CLIP:层级条件双曲视觉-语言训练

Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat, Seifeldin Elkerdany, Weixing Wang, Gerard de Melo

机构 * hyper 3 labs(超3实验室) Hasso Plattner Institute(哈索·普拉特纳研究所) University of Potsdam(波茨坦大学) Faculty of Computers and Data Science, Alexandria University(亚历山大大学计算机与数据科学学院) Faculty of Computer Science and Engineering, Alamein International University(阿拉曼国际大学计算机科学与工程学院)

专题命中 其他LLM :language model(abstract);分类 cs.AI

AI总结 Hyper3-CLIP结合层级条件与双曲几何,通过文本构建多粒度查询层级训练视觉-语言模型,提升图像-文本检索与多标签分类性能,在层级指标上保持竞争力。

Comments 16 pages, 2 figures, ECCV 2026 Beyond Euclidean Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29133 2026-09-01 cs.CL 新提交 57%

AI Historian: Helping historians organize and verify person-centred temporal clues from dispersed historical narratives

AI历史学家:帮助历史学家从分散的历史叙事中整理和验证以人为中心的时间线索

Yifeng Lu, Zijie Yang, Jie Li, Qingkai Min, Yue Zhang

机构 * Westlake University(西湖大学) Peking University(北京大学)

专题命中 其他LLM :prompting(abstract);分类 cs.CL

AI总结 研究提出AIH智能体系统,可从分散历史叙事中整理验证时间线索,在《史记》案例中表现优于人类及大语言模型,已应用于多国历史材料并降低整理成本。

Comments 37 pages, 21 figures. Code: https://github.com/YiFengLu1999/AI-Historian

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27860 2026-08-31 cs.CV cs.AI 新提交 57%

From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

从透视图像到鱼眼图像的深度估计与开放词汇分割

Rit Gangopadhyay, Alex Wong

机构 * Yale University(耶鲁大学)

专题命中 其他LLM :foundation model(abstract);分类 cs.AI

AI总结 该研究提出了与架构、任务无关的Distortion Extenders(DEX),通过自监督对齐损失将视觉基础模型泛化到鱼眼相机,在深度估计和开放词汇分割任务中优于基线,还可用于相机校准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27843 2026-08-31 cs.CL cs.MA 新提交 57%

Synthetic Linguistic Agency: How an Embodied Mortal Agent Learns Linguistic Affordances through Consequential Social Experience

合成语言智能体:具身化有限生命智能体如何通过具有因果关系的社会经验学习语言可供性

Sixin Chen, Taizhou Chen

机构 * College of Engineering, Shantou University(汕头大学工程学院) Department of Computer Science, Shantou University(汕头大学计算机科学系)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 本研究提出合成语言智能体(SLA)的可检验标准,开发具身化有限生命智能体(EMA)模型,证实其能通过社会经验学习语言可供性并展现SLA,为合成共情与人机交互研究提供支撑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27760 2026-08-31 cs.CL 新提交 57%

Informational Antilocality and the Locality Bias in LLMs

信息非局部性与大型语言模型(LLMs)的局部性偏差

Andrew McInnerney, Shane Storks, Steven Abney, Richard L. Lewis

机构 * University of Michigan(密歇根大学) Eastern Michigan University(东密歇根大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 该研究探讨LLMs学习k-非局部语言的能力,发现其在不同k值非局部语言上交叉熵损失相当,但非局部性越强收敛越慢,为非局部依赖更难学习的观点提供了学习速度层面的证据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16574 2026-08-31 cs.CL 版本更新 57%

Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

Transformer预测在人类句子处理中的分歧:对一致吸引效应的全面分析

Titus von der Malsburg, Sebastian Padó

机构 * University of Stuttgart(斯图加特大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 本文通过对比不同大小和架构的Transformer模型,评估其在英语一致吸引配置中的表现,发现模型在处理宾语提取相对从句时表现不佳,无法复制人类的不对称干扰模式。

Comments Paper accepted for EMNLP 2026 main conference

Journal ref Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17654 2026-08-31 cs.IR cs.LG 版本更新 57%

Mine and Refine: Optimizing Graded Relevance in E-commerce Semantic Search Retrieval

挖掘与精炼:优化电子商务搜索检索中的分级相关性

Jiaqi Xi, Raghav Saboo, Luming Chen, Johny Rufus, Aditya Dodda, Ved Sampath, Kenny Chi, Elyse Winer, Akshad Viswanathan, Martin Wang, Sudeep Das

机构 * DoorDash Inc.(DoorDash公司)

专题命中 其他LLM :LLM(abstract);分类 cs.LG

AI总结 本文提出一种两阶段框架,通过挖掘和精炼提升电子商务搜索检索的分级相关性,结合对比学习和多类圆损失优化嵌入空间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19449 2026-08-28 cs.SE cs.LG 版本更新 57%

Stack Trace-Based Crash Deduplication with Transformer Adaptation

基于栈跟踪的Transformer适配崩溃重复检测

Md Afif Al Mamun, Gias Uddin, Lan Xia, Longyu Zhang

机构 * University of Calgary(卡尔加里大学) York University(约克大学) IBM Canada(IBM加拿大)

专题命中 其他LLM :language model(abstract);分类 cs.LG

AI总结 提出基于Transformer的dedupT方法,适配预训练语言模型处理栈跟踪,训练全连接网络排序重复崩溃,在四个公开数据集上优于现有方法,减少人工分类工作量。

Comments This work is currently under review at IEEE Transactions on Software Engineering (TSE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25505 2026-08-27 cs.IT cs.CL math.IT 新提交 57%

Conditional Total Correlation and the Serial Depth of Adaptive Parallel Sampling

条件总相关与自适应并行采样的序列深度

Chuling Wen, Weijie Liang, Jian Lu

机构 * Shenzhen University(深圳大学) National Center for Applied Mathematics Shenzhen (NCAMS)(深圳国家应用数学中心)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 该研究针对离散向量自适应并行采样,通过推导条件总相关与采样散度的恒等式,明确序列深度的决定因素,并在马尔可夫链等场景验证其特性,为掩码扩散模型解码提供理论支撑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18455 2026-08-27 cs.CY cs.AI 版本更新 57%

Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia

AI搜索摘要对网站流量的影响:来自谷歌AI概览和维基百科的证据

Mehrzad Khosravi, Hema Yoganarasimhan

机构 * University of Washington(华盛顿大学)

专题命中 其他LLM :LLM(abstract_cn);分类 cs.AI

AI总结 研究通过对比谷歌AI概览与维基百科多语言版本流量变化,发现AI摘要导致英文维基百科文章日均流量下降15%,且对文化类内容影响更大,揭示搜索引擎生成答案对信息类出版物注意力的再分配。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23570 2026-08-26 cs.CL 新提交 57%

Taming Visual Neglect: A Variational Information Bottleneck Framework for Adaptive Attention in Multimodal In-Context Learning

驯服视觉忽视:用于多模态上下文学习中自适应注意力的变分信息瓶颈框架

Kaito Tanaka, Yuji Nishimura, Keisuke Matsuda, Aya Nakayama

机构 * SANNO University(山王大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 针对多模态上下文学习中视觉上下文时而被利用时而被忽视的问题,提出VIB-ICL框架,通过CMIG量化跨模态信息,推导理论界并经五组基准实验验证,实现准确率提升与演示样本减少。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13092 2026-08-26 cs.LG cs.AR 版本更新 57%

Breaking the Tuning Barrier: Zero-Hyperparameters Yield Multi-Corner Analysis Via Learned Priors

打破调优壁垒:通过先验学习实现零超参数的多角分析

Wei W. Xing, Kaiqi Huang, Jiazhan Liu, Hong Qiu, Shan Shen

机构 * School of Mathematical and Physical Science, University of Sheffield(谢菲尔德大学数学与物理科学学院) SZU–UoS Joint Centre for Innovation and Entrepreneurship, College of Mechatronics and Control Engineering, Shenzhen University(深大-乌兹别克斯坦联合创新与创业中心,机电控制工程学院,深圳大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 其他LLM :foundation model(abstract);分类 cs.LG

AI总结 针对电路多角分析中仿真成本高且现有方法需大量调参的问题,提出基于预训练基础模型的上下文学习方法,无需调优即可匹配最先进精度,将验证成本降低10倍以上。

Comments Published in DAC2026. Final Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05387 2026-08-24 cs.CL 57%

The Generalization Ridge: Information Flow in Natural Language Generation

通用 ridge:自然语言生成中的信息流

Ruidi Chang, Chunyuan Deng, Hanjie Chen

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 研究揭示了 transformer 模型中信息流的非单调趋势,通过信息理论框架分析隐藏表示与目标输出间的互信息变化,发现中间层形成泛化 ridge,为理解模型泛化机制提供新视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19430 2026-08-21 cs.IR cs.CL cs.CR 新提交 57%

HARP: Hierarchical Adaptive Ranking with Preference-Adaptive Fusion for Query-Based CVE Prioritization

HARP:用于基于查询的CVE优先级排序的分层自适应偏好融合排名方法

Haochen Liu, Zhengzhang Chen, Haoyu Wang, Yanchi Liu, Jundong Li, Haifeng Chen

机构 * University of Virginia(弗吉尼亚大学) NEC Laboratories America(美国 NEC 实验室)

专题命中 其他LLM :LLM(abstract_cn);分类 cs.CL

AI总结 本文针对基于查询的CVE优先级排序问题,提出HARP框架,结合漏洞知识图谱与历史标注示例,在多场景下优于多个基线方法,实现了更有效的优先级排序。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19369 2026-08-21 cs.CL cs.CR math.DG 新提交 57%

Linguistic Holonomy and Statistical Watermarks: Inner Geometry of Meaning-Preserving Transformations

语言完整性与统计水印:保留意义变换的内蕴几何

Daniele Corradetti

机构 * Instituto Superior Técnico(高等技术学院) Universidade do Algarve(阿尔加维大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 该研究针对语言模型统计水印的失效问题,借助语言环形式体系与威尔逊环类比,推导得出残差统计量与种子窗口存活位置数的精确恒等式,揭示相同保留率下信号存活量随编辑位置变化的规律。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18041 2026-08-21 cs.CL 版本更新 57%

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

语言具有两个参数:叙事诱导的语义可塑性与相位敏感的解读

Hollis Robbins

机构 * University of Utah(犹他大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 该研究提出语言解读需第二个带符号、持久且以个体与二元组为索引的相位参数,指出标准Transformer无相位显式表征,需构建承载相位语义状态的语言模型。

Comments 23 pages; 0 figuresCC

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18167 2026-08-20 cs.AI cs.SE 新提交 57%

Adversarial Review: Structured Disagreement for Grounded Agentic Code Review

对抗评审:基于结构化分歧的 grounded 智能体代码评审

Eric S. Qiu, Joyce Gill

机构 * Cornell University(康奈尔大学) Stanford University(斯坦福大学)

专题命中 其他LLM :LLM(abstract);分类 cs.AI

AI总结 本文提出对抗评审(AR)协议,仅用3个智能体实现结构化分歧的协作代码评审,在LiveCodeBench、SWE-PRBench等基准上优于多智能体基线,证明无需大量智能体即可完成高效代码评审。

Comments Accepted to ICML 2026 Workshop on DL4C

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17202 2026-08-19 cs.AI cs.CR 新提交 57%

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

愚人金:针对开放权重模型安全移除攻击的防御性欺骗

Mark Russinovich

机构 * Microsoft Azure(微软Azure)

专题命中 其他LLM :language model(abstract);分类 cs.AI

AI总结 针对开放权重模型易被移除安全对齐的问题,提出“愚人金”防御,通过训练仅在被攻击状态显现的伪造诱饵,使被攻击模型对危险请求输出高比例假回答,且无法区分真假,同时不影响干净状态的良性行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00618 2026-08-19 cs.CV cs.AI 57%

DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition

DesCLIP:通过通用属性描述实现基于视觉语言模型的鲁棒持续学习

Chiyuan He, Zihuan Qiu, Fanman Meng, Linfeng Xu, Qingbo Wu, Hongliang Li

机构 * School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院)

专题命中 其他LLM :language model(abstract);分类 cs.AI

AI总结 本文提出DesCLIP,通过通用属性描述引导视觉语言模型理解特定类别对象,建立视觉-通用属性-类别三元关联,提升持续学习效果。

Comments IEEE Transactions on Multimedia 2026

Journal ref IEEE Transactions on Multimedia, vol. 28, pp. 5021-5035, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06505 2026-08-19 cs.CL 版本更新 57%

Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning

在语境中说话:通过对比学习实现的多语言语音识别与语音语境对齐

Yuchen Zhang, Haralambos Mouratidis, Ravi Shekhar

机构 * Institute for Analytics and Data Science, University of Essex(数据分析与科学研究所,埃塞克斯大学) School of Computer Science and Electronic Engineering, University of Essex(计算机科学与电子工程学院,埃塞克斯大学)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 本文提出一种基于对比学习的多语言语音识别框架,通过语境对齐提升识别性能,支持多种语言和口音,实现语音与上下文的高效交互。

Comments Accepted at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15910 2026-08-18 eess.AS cs.CL cs.SD 新提交 57%

Iterative Self-Learning for Expressive Text-to-Speech Synthesis

用于高表现力文本到语音合成的迭代自学习

Nicholas Sanders, Gustav Eje Henter, Simon King, Korin Richmond

机构 * University of Edinburgh(爱丁堡大学) Huawei(华为) Wallenberg AI, Autonomous Systems and Software Program(瓦伦堡人工智能、自主系统与软件计划)

专题命中 其他LLM :prompting(abstract);分类 cs.CL

AI总结 针对低资源下高表现力TTS的标签稀缺问题,提出迭代自学习框架,基于Invert-Classify方法迭代优化伪标签,在词级突出性和情感任务上验证,提升了标签贴合度与合成质量,性能接近全监督模型。

详情

展开后加载摘要…

URL PDF HTML 收藏