arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-15 至 2026-01-15 共收录 9 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 9 篇

2601.08843 2026-01-15 cs.CL cs.LG 86%

Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness

基于评分标准的LLM评分:对齐、不确定性与鲁棒性

Haotian Deng, Chris Farber, Jiyoon Lee, David Tang

机构 * Purdue University(普渡大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文研究了基于评分标准的LLM评分的对齐性、不确定性与鲁棒性,发现评分标准粒度增加时对齐性下降,通过信任曲线分析发现过滤低置信度预测可提升准确性,同时模型对同义词替换敏感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08415 2026-01-15 cs.CY cs.AI 85%

Regulatory gray areas of LLM Terms

大型语言模型条款的监管灰色区域

Brittany I. Davidson, Kate Muir, Florian A. D. Burnat, Adam N. Joinson

机构 * University of Bath(巴斯大学) Bath Spa University(巴斯斯巴大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文分析了五家主要LLM提供商的条款,揭示了使用限制的差异及监管灰色区域,提供了公开资源以帮助用户和研究人员应对这一挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23335 2026-01-15 cs.CL cs.AI cs.DB 81%

Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach

通过结构化知识发现方法提升语言模型生成的可解释性

Shuqi Liu, Han Wu, Guanzhi Deng, Jianshu Chen, Xiaoyang Wang, Linqi Song

机构 * City University of Hong Kong(香港城市大学) City University of Hong Kong Shenzhen Research Institute(香港城市大学深圳研究院) Tencent AI Lab(腾讯AI实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种任务无关的结构化知识猎手,通过结合语言模型的生成能力与知识猎手的高保真度,提升语言模型生成文本的可解释性,并在多个数据集上验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11311 2026-01-15 eess.IV cs.AI cs.CV cs.LG 81%

Large-scale modality-invariant foundation models for brain MRI analysis: Application to lesion segmentation

大规模模态不变基础模型用于脑MRI分析:应用于病变分割

Petros Koutsouvelis, Matej Gazda, Leroy Volmer, Sina Amirrajab, Kamil Barbierik, Branislav Setlak, Jakub Gazda, Peter Drotar

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种大规模模态不变基础模型,用于提升脑MRI中病变分割的性能,通过自监督学习预训练并保留细粒度模态特定特征。

Comments Submitted to IEEE ISBI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08860 2026-01-15 cs.CV cs.AI 79%

Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models

视觉-语言模型和生成图像模型中的偏见检测与旋转鲁棒性缓解

Tarannum Mithila

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文提出旋转鲁棒缓解策略,通过数据增强、表征对齐和模型正则化,提升视觉-语言和生成图像模型在旋转和分布偏移下的鲁棒性和公平性。

Comments Preprint. This work is derived from the author's Master's research. Code and supplementary materials will be released separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08876 2026-01-15 cs.CV 78%

The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models

具身AI中的语义生命周期:通过基础模型实现获取、表示与存储

Shuai Chen, Hao Chen, Yuanchen Bei, Tianyang Zhao, Zhibo Zhou, Feiran Huang

机构 * College of Information Science and Technology, Jinan University(信息科学与技术学院,暨南大学) Faculty of Data Science, City University of Macau(数据科学学院,澳门城市大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Zhongguancun Laboratory(中关村实验室) College of Cyber Security, Jinan University(网络安全学院,暨南大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本文提出语义生命周期框架,通过基础模型在具身AI中实现语义信息的获取、表示与存储,探讨了语义处理的连续流动与维护,并总结了当前挑战与未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21364 2026-01-15 cs.LG cs.AI 73%

Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders

无需牺牲的可解释性:混合解码器的忠实密集层分解

James Oldfield, Shawn Im, Sharon Li, Mihalis A. Nicolaou, Ioannis Patras, Grigorios G Chrysos

机构 * UW-Madison(威斯康星大学麦迪逊分校)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出混合解码器(MxDs)通过层级稀疏性实现密集层分解,保持原始解码器的表达能力,显著提升稀疏性-准确性前沿性能。

Comments NeurIPS 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21518 2026-01-15 cs.CV cs.CL cs.LG 62%

Head Pursuit: Probing Attention Specialization in Multimodal Transformers

头部追踪:探究多模态转换器中的注意力专业化

Lorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello, Alberto Cazzaniga

机构 * Area Science Park(面积科学公园) Sapienza University of Rome(罗马萨皮恩扎大学) Institute of Science and Technology(科学与技术研究所)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.LG

AI总结 本研究通过分析多模态转换器中注意力头的专业化,揭示了模型内部可控的结构,并展示了通过编辑少量头部以增强或抑制特定概念的可行性。

Comments Accepted at NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03997 2026-01-15 cs.LG cs.CL 62%

Quiet Feature Learning in Algorithmic Tasks

算法任务中的安静特征学习

Prudhviraj Naidu, Zixian Wang, Leon Bergen, Ramamohan Paturi

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.LG

AI总结 该研究发现,在算法任务中,模型在损失曲线平坦区域学习了对任务性能至关重要的中间特征,挑战了传统损失作为学习代理的假设。

Comments Accepted as Oral presentation @ AAAI 2026 Special Track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏