arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7565 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7565 篇

2601.13886 2026-01-21 cs.CV 67%

Revisiting Multi-Task Visual Representation Learning

重新审视多任务视觉表示学习

Shangzhe Di, Zhonghua Zhai, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学SAI) ByteDance Seed(字节跳动种子)

专题命中 知识编辑与模型理解 :language model(abstract);pretraining(abstract)

AI总结 MTV框架通过多任务学习结合视觉-语言对比、自监督和密集空间目标,提升空间推理能力的同时保持全局语义理解。

Comments Code: https://github.com/Becomebright/MTV

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12303 2026-01-21 cs.CV 67%

Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations

表示法中的概念:通过视觉表示的稀疏分解实现事后概念瓶颈模型

Shizhan Gong, Xiaofan Zhang, Qi Dou

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文提出PCBM-ReD,通过视觉表示的稀疏分解,实现对预训练模型的可解释性增强,提升图像分类任务的准确性和可解释性。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11182 2026-01-19 cs.IR 67%

From Knots to Knobs: Towards Steerable Collaborative Filtering Using Sparse Autoencoders

从结到旋钮:迈向可操控的协作过滤使用稀疏自编码器

Martin Spišák, Ladislav Peška, Petr Škoda, Vojtěch Vančura, Rodrigo Alves

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文首次将稀疏自编码器应用于协作过滤,通过在编码器和解码器之间插入SAE来增强协作自编码器,以实现推荐方向的可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05192 2026-01-19 cs.SE 67%

AI-assisted JSON Schema Creation and Mapping

基于AI的JSON模式创建与映射

Felix Neubauer, Jürgen Pleiss, Benjamin Uekermann

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文提出结合大型语言模型与确定性技术的混合方法,用于基于自然语言输入的JSON模式创建、修改和映射,降低非专家在结构化数据建模与集成中的门槛。

Journal ref 2025 ACM/IEEE 28th International Conference on Model Driven Engineering Languages and Systems Companion (MODELS-C), Grand Rapids, MI, USA, 2025, pp. 79-83

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06224 2026-01-14 cs.CV 67%

Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization

在所见之地接地:通过标题反馈、多样性感知采样和冲突正则化实现抗幻觉的MLLMs

Miao Pan, Wangjie Gan, Jintao Chen, Wenqi Zhang, Bing Sun, Jianwei Yin, Xuhong Zhang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文提出抗幻觉的MLLMs方法,通过标题反馈、多样性感知采样和冲突正则化减少幻觉,提升推理准确性。

Comments AAAI-2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14367 2026-01-12 cs.CV 67%

Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution

幻觉分数:朝着减轻生成图像超分辨率中的幻觉

Weiming Ren, Raghav Goyal, Zhiming Hu, Tristan Ty Aumentado-Armstrong, Iqbal Mohomed, Alex Levinshtein

机构 * University of Waterloo(多伦多大学) AI Center – Toronto, Samsung Electronics(三星电子人工智能中心)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文提出幻觉分数以衡量和减轻生成图像超分辨率中的幻觉问题,通过多模态大语言模型生成 HS 并用于模型微调。

Comments 31 pages, 21 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16996 2026-01-12 econ.TH cs.GT 67%

Artificial Intelligence Clones

人工智能克隆

Annie Liang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文研究了人工智能克隆在匹配搜索中的影响,发现有限的面对面接触比无限的人工智能表示搜索更有效,尤其是在高维人格空间中。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13729 2026-01-06 cs.CY cs.GT 67%

The Economics of Information Pollution in the Age of AI: General Equilibrium, Welfare, and Policy Design

人工智能时代的信息污染经济学:一般均衡、福利与政策设计

Yukun Zhang, Tianyang Zhang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文研究人工智能时代信息污染的经济学问题,通过一般均衡模型分析信息污染的市场失灵,并提出基于信息污染指数的适应性治理框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21004 2025-12-25 cs.CV 67%

Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations

从下一帧预测学习:自回归视频建模编码有效表示

Jinghan Li, Yang Jin, Hao Jiang, Yadong Mu, Yang Song, Kun Xu

机构 * Peking University(北京大学)

专题命中 知识编辑与模型理解 :foundation model(abstract);pretraining(abstract)

AI总结 NExT-Vid通过掩码下一帧预测提出自回归视频预训练框架,提升视频生成质量和语义表示能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20034 2025-12-24 cs.IR 67%

VSA:Visual-Structural Alignment for UI-to-Code

VSA:面向UI到代码的视觉-结构对齐

Xian Wu, Ming Zhang, Zhiyu Fang, Fei Li, Bin Wang, Yong Jiang, Hao Zhou

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 VSA通过视觉-结构对齐提升UI到代码的模块化和一致性,生成类型安全的组件以提高软件工程效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18671 2025-12-23 cs.CV 67%

SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse

SmartSight: 通过时间注意力崩溃缓解视频大语言模型中的幻觉而不影响视频理解

Yiming Sun, Mi Zhang, Feifei Li, Geng Hong, Min Yang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 SmartSight通过时间注意力崩溃技术,在不牺牲视频理解能力的前提下,有效降低视频大语言模型的幻觉问题。

Comments AAAI26 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15068 2025-12-22 cs.LG cs.AI cs.CL 67%

The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems

语义幻觉:基于嵌入的幻觉检测在RAG系统中的认证极限

Debu Sinha

机构 * Independent Researcher(独立研究者)

专题命中 知识编辑与模型理解 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究揭示了基于嵌入的幻觉检测在RAG系统中存在语义幻觉问题,通过符合预测方法发现真实幻觉检测的挑战,证明需通过推理而非表面语义解决。

Comments 12 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15410 2025-12-18 cs.CV 67%

Preserving Marker Specificity with Lightweight Channel-Independent Representation Learning

通过轻量级通道无关表征学习保持标记特异性

Simon Gutwein, Arthur Longuefosse, Jun Seita, Sabine Taschner-Mandl, Roxane Licandro

机构 * St. Anna Children’s Cancer Research Institute(圣安娜儿童癌症研究中心) TU Wien, Institute of Visual Computing and Human-Centered Technology, CVL(维也纳技术大学,视觉计算与人本技术研究所,CVL) Medical University of Vienna, Biomedical Imaging and Image-guided Therapy, Computational Imaging Research, ELIA Group(维也纳医学大学,生物医学成像与影像引导治疗,计算成像研究,ELIA小组) RIKEN Center for Integrative Medical Sciences(理化学研究所整合医学中心) Medical University of Vienna, Comprehensive Center for AI in Medicine(维也纳医学大学,医学人工智能综合中心) Medical University of Vienna, Christian Doppler Lab for Mathematical Modelling and Simulation of Next-Generation Medical Ultrasound Devices(维也纳医学大学,基督教多普勒实验室(下一代医学超声设备数学建模与模拟))

专题命中 知识编辑与模型理解 :foundation model(abstract);pretraining(abstract)

AI总结 本文提出轻量级通道无关模型CIM-S,通过保持标记独立性和浅层架构,在多路数据自监督学习中优于传统深度模型。

Comments 16 pages, 9 figures, MIDL 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05745 2025-12-08 cs.CR cs.MM 67%

ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior

ARGUS: 通过引导指令遵循行为防御多模态间接提示注入攻击

Weikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li, Huiping Zhuang, Ruidong Wang, Cen Chen, Hao Peng

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 ARGUS通过引导指令遵循行为,在表示空间中寻找最优防御方向,实现对多模态间接提示注入攻击的有效防御,同时保持模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02981 2025-12-03 cs.CV 67%

InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration

InEx:通过内省与跨模态多智能体协作缓解幻觉

Zhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An, Wei Pang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 InEx通过内省推理和跨模态多智能体协作,自主缓解大型语言模型的幻觉问题,实验表明其在多个基准上表现优异。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20273 2025-11-26 cs.LG cs.AI cs.CL 67%

Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits

超越组件:基于奇异向量的Transformer电路可解释性

Areeb Ahmad, Abhinav Joshi, Ashutosh Modi

机构 * Indian Institute of Technology Kanpur (IIT Kanpur)(印度理工学院坎浦尔(IIT坎浦尔))

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于奇异向量的Transformer电路可解释性方法,揭示了模型内部更分布、结构化和组合化的计算特性。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17645 2025-11-26 cs.LG cs.AI cs.CL 67%

BlockCert: Certified Blockwise Extraction of Transformer Mechanisms

BlockCert: Transformer机制的认证分块提取

Sandro Andric

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 BlockCert通过认证分块提取Transformer机制,提供可验证的误差界和覆盖率指标,实验证明其在多个模型上有效且具有实际应用价值。

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15214 2025-11-26 q-fin.GN cs.CY 67%

Corporate Earnings Calls and Analyst Beliefs

企业盈利电话会议与分析师信念

Giuseppe Matera

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 研究通过分析企业盈利电话会议中的叙述,发现分析师对乐观情绪过度反应,对风险和不确定性的叙述反应不足,揭示了预期形成的机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17481 2025-11-24 cs.CV 67%

Counterfactual World Models via Digital Twin-conditioned Video Diffusion

通过数字孪生条件的视频扩散实现反事实世界模型

Yiqing Shen, Aiza Maksutova, Chenjia Li, Mathias Unberath

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 CWMDT通过构建数字孪生和应用大语言模型,实现反事实世界模型,提升对视频正向模拟的控制能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16769 2025-11-24 cs.HC 67%

Trust in AI emerges from distrust in humans: A machine learning study on decision-making guidance

对人工智能的信任源于对人类的不信任:一项关于决策指导的机器学习研究

Johan Sebastián Galindez-Acosta, Juan José Giraldo-Huertas

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本研究通过机器学习探讨了人工智能在决策指导中的信任机制,发现对人类的不信任会促使人们转向AI,且AI在事实性情景中更受青睐。

Comments 36 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03725 2025-11-24 cs.CV 67%

Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition

解构概念胜于言语:可解释的视频动作识别

Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim, Jinwoo Choi

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 DANCE框架通过分离运动动态、物体和场景的概念类型,提升视频动作识别的可解释性与性能。

Comments NeurIPS 2025 Spotlight paper. Project page: https://jong980812.github.io/DANCE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16435 2025-11-21 cs.CV 67%

Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation

超越视觉线索:利用通用语义作为少样本分割的支持

Jin Wang, Bingfeng Zhang, Jian Pang, Mengyu Liu, Honglong Chen, Weifeng Liu

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文提出语言驱动属性泛化架构,通过多属性增强和多模态对齐提升少样本分割性能,实现新的最佳效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02091 2025-11-19 cs.CV 67%

MoReFun: Past-Movement Guided Motion Representation Learning for Future Motion Prediction and Understanding

Junyu Shi, Haoting Wu, Zhiyuan Zhang, Lijiang Liu, Yong Sun, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 知识编辑与模型理解 :LLM(abstract);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10059 2025-11-14 cs.CV 67%

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, Yu Zhou

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19361 2025-11-14 cs.CV 67%

ImageSet2Text: Describing Sets of Images through Text

Piera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia, Nuria Oliver

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08641 2025-11-13 cs.CR cs.CY cs.MA 67%

QOC DAO -- Stepwise Development Towards an AI Driven Decentralized Autonomous Organization

Marc Jansen, Christophe Verdot

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21412 2025-11-11 cs.CV 67%

Bridging the gap to real-world language-grounded visual concept learning

Whie Jung, Semin Kim, Junee Kim, Seunghoon Hong

机构 * School of Computing, KAIST(计算机学院,韩国科学技术院)

专题命中 知识编辑与模型理解 :language model(abstract);prompting(abstract)

Journal ref Advances in Neural Information Processing Systems (NeurIPS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11295 2025-11-03 cs.CV 67%

Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering

Jian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao, Yihong Liu, Hinrich Schütze, Michael A. Hedderich, Thomas Seidl

机构 * University of Munich(慕尼黑大学) Munich Center of Machine Learning(慕尼黑机器学习中心)

专题命中 知识编辑与模型理解 :language model(abstract);SFT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23953 2025-11-03 cs.LG cs.AI cs.CL cs.CY cs.GT 67%

Representative Social Choice: From Learning Theory to AI Alignment

Tianyi Qiu

机构 * Peking University(北京大学) UC Berkeley(加州大学伯克利分校) Center for Human-Compatible AI(人类兼容人工智能中心)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Journal of Artificial Intelligence Research, in press. Best Paper at NeurIPS 2024 Pluralistic Alignment Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25798 2025-10-31 cs.LG cs.AI cs.CL 67%

MemEIC: A Step Toward Continual and Compositional Knowledge Editing

Jin Seong, Jiyun Park, Wencke Liermann, Hongseok Choi, Yoonji Nam, Hyun Kim, Soojong Lim, Namhoon Lee

机构 * Electronics and Telecommunications Research Institute, Republic of Korea(韩国电子电信研究院) POSTECH Sungkyunkwan University(全南大学)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025, 38 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏