arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7552 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7552 篇

2506.23951 2026-02-23 cs.CL 70%

Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders

揭示文本分类中大语言模型的决策过程:使用稀疏自编码器提取影响性且可解释的概念

Mathis Le Bail, Jérémie Dentan, Davide Buscaldi, Sonia Vanier

机构 * LIX (École Polytechnique, IP Paris, CNRS)(LIX(巴黎理工学院,IP巴黎,CNRS)) LIPN (Sorbonne Paris Nord)(LIPN(巴黎-索邦大学北校区))

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出ClassifSAE模型,通过稀疏自编码器提升文本分类中特征的因果性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18145 2026-02-23 cs.CL 70%

Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention

基于频率感知注意力的LLM幻觉检测

Siya Qi, Yudong Chen, Runcong Zhao, Qinglin Zhu, Zhanghao Hu, Wei Liu, Yulan He, Zheng Yuan, Lin Gui

机构 * Department of Informatics, King's College London, UK(伦敦国王学院信息学院) Department of Statistics, University of Warwick, UK(沃里克大学统计系) School of Computer Science, The University of Sheffield, UK(谢菲尔德大学计算机科学学院) The Alan Turing Institute, UK(艾伦·图灵研究所)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出一种基于频率感知注意力的轻量级幻觉检测方法,通过分析生成过程中注意力的高频成分,有效识别幻觉token,提升了LLM在上下文生成中的可靠性。

Comments 25 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15871 2026-02-19 cs.CL cs.CY 70%

CheckIfExist: Detecting Citation Hallucinations in the Era of AI-Generated Content

CheckIfExist:在AI生成内容时代检测引用幻觉

Diletta Abbonato

机构 * Unito(乌托大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 CheckIfExist是一个开源工具,通过多源验证检测AI生成内容中的引用幻觉,提供即时验证和批量处理功能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00364 2026-02-19 cs.CV cs.LG 70%

LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation

LMSeg:释放大规模模型在开放词汇语义分割中的潜力

Huadong Tang, Youpeng Zhao, Yan Huang, Min Xu, Jun Wang, Qiang Wu

机构 * Huadong Tang(华东唐) Youpeng Zhao(赵有鹏) Yan Huang(黄艳) Min Xu(徐敏) Jun Wang(王俊) Qiang Wu(吴强)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 LMSeg通过结合多个大规模模型提升开放词汇语义分割的性能,利用LLMs生成多样化视觉属性的提示并改进视觉特征提取。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15509 2026-02-18 cs.CL 70%

Fine-Refine: Iterative Fine-grained Refinement for Mitigating Dialogue Hallucination

细粒度细化:用于缓解对话幻觉的迭代细化

Xiangyan Chen, Yujian Gan, Matthew Purver

机构 * Queen Mary University of London(伦敦女王学院) Queen’s University Belfast(贝尔法斯特女王大学) Institut Jožef Stefan(乔泽夫·斯蒂芬研究所)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 Fine-Refine通过细粒度细化框架提升对话系统事实准确性,实现7.63分的提升,同时牺牲少量对话质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14529 2026-02-17 cs.AI 70%

Disentangling Deception and Hallucination Failures in LLMs

解构大语言模型中的欺骗与幻觉故障

Haolang Lu, Hongrui Peng, WeiYe Fu, Guoshun Nan, Xinye Cao, Xingrui Li, Hongcan Guo, Kun Wang

机构 * Beijing University of Posts(北京邮电大学) Nanyang Technological University(南洋理工大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究提出一种机制导向的视角,解构大语言模型中幻觉与欺骗的不同故障机制,通过受控环境分析四种行为案例,揭示知识存在与行为表达的分离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12714 2026-02-16 cs.LG 70%

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools -- From Consensus Learning to Ambiguity-Driven Emotion Reasoning

ADEPT: 通过证据探测工具实现情感的代理解码 -- 从共识学习到由模糊性驱动的情感推理

Esther Sun, Bo-Hao Su, Abinay Reddy Naini, Shinji Watanabe, Carlos Busso

机构 * cmu(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 ADEPT通过证据探测工具实现情感的代理解码,从共识学习转向由模糊性驱动的情感推理,提升情感识别的准确性和可解释性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12058 2026-02-13 cs.SE cs.AI cs.FL 70%

ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair

ModelWisdom: 一个集成工具包用于TLA+模型可视化、摘要与修复

Zhiyong Chen, Jialun Cao, Chang Xu, Shing-Chi Cheung

机构 * State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家重点实验室,南京大学,中国) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学计算机科学与工程系)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ModelWisdom通过可视化、图优化、摘要和修复功能,提升TLA+模型检查的可解释性和可操作性。

Comments Accepted by FM 2026 Research Track (Tool)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10504 2026-02-12 cs.CL 70%

On the Robustness of Knowledge Editing for Detoxification

关于知识编辑在去毒化中的鲁棒性

Ming Dong, Shiyi Tang, Ziyan Peng, Guanyi Chen, Tingting He

机构 * Hubei Provincial Key Laboratory of Artificial Intelligence(人工智能与智能学习湖北省重点实验室) Research Center for Network Media, School of Computer Science, Central China Normal University, Wuhan, China(网络媒体研究中心,计算机科学学院,中央民族大学,武汉,中国)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出了一种评估基于知识编辑的去毒化方法鲁棒性的框架,发现其效果受模型、目标和语言的限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21767 2026-02-12 cs.CL 70%

Evaluating ChatGPT on Medical Information Extraction Tasks: Performance, Explainability and Beyond

评估ChatGPT在医学信息抽取任务中的表现:性能、可解释性及更进一步

Liz Li, Wei Zhu

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 评估ChatGPT在医学信息抽取任务中的性能、可解释性及应用挑战

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25069 2026-02-11 cs.CL 70%

TOPol: Capturing and Explaining Multidimensional Semantic Polarity Fields and Vectors

TOPol: 捕捉和解释多维语义极性场和向量

Gabin Taibi, Lucia Gomez

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 TOPol通过半无监督框架捕捉和解释多维语义极性场,用于上下文敏感的多维叙述分析。

Comments 7 pages, 3 figures and 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22730 2026-02-10 cs.CV cs.AI eess.IV 70%

Improved cystic hygroma detection from prenatal imaging using ultrasound-specific self-supervised representation learning

改进的产前影像中囊性水瘤检测:利用超声特定的自监督表示学习

Youssef Megahed, Robin Ducharme, Inok Lee, Inbal Willner, Adrian D. C. Chan, Mark Walker, Steven Hawken

机构 * organization= Department of Systems Computer Engineering, Carleton University , city= Ottawa , state= Ontario , country= Canada organization= Department of Methodological Implementation Research, Ottawa Hospital Research Institute , city= Ottawa , state= Ontario , country= Canada organization= Department of Acute Care Research, Ottawa Hospital Research Institute , city= Ottawa , state= Ontario , country= Canada organization= Children's Hospital of Eastern Ontario Research Institute , city= Ottawa , state= Ontario , country= Canada organization= Better Outcomes Registry \& Network Ontario, Children’s Hospital of Eastern , city= Ottawa , state= Ontario , country= Canada organization= Department of Obstetrics Gynecology, University of Ottawa , city= Ottawa , state= Ontario , country= Canada organization= School of Epidemiology Public Health, University of Ottawa , city= Ottawa , state= Ontario , country= Canada organization= Department of Obstetrics, Gynecology \& Newborn Care, The Ottawa Hospital , city= Ottawa , state= Ontario , country= Canada Global Health Office, University of Ottawa , city= Ottawa , state= Ontario , country= Canada organization= Department of Clinical Science Translational Medicine, University of Ottawa , city= Ottawa , state= Ontario , country= Canada

专题命中 知识编辑与模型理解 :foundation model(abstract);pretraining(abstract);分类 cs.AI

AI总结 本研究提出利用超声特定自监督学习方法,改进产前超声图像中囊性水瘤的检测,显著提升准确率和AUC指标。

Comments 13 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15735 2026-02-09 cs.LG 70%

EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs

EigenTrack:基于隐藏激活谱几何的特征追踪用于LLMs和VLMs中的幻觉和分布外检测

Davide Ettori, Nastaran Darabi, Sina Tayebati, Ranganath Krishnan, Mahesh Subedar, Omesh Tickoo, Amit Ranjan Trivedi

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Capital One AI Labs(Capital One AI实验室) Intel Labs(英特尔实验室)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 EigenTrack通过分析隐藏激活的谱几何,提供一种实时检测LLMs和VLMs中幻觉和分布外错误的可解释方法。

Comments 5 pages, submitted to ICASSP 2026, September 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05927 2026-02-06 stat.ML cs.LG 70%

Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences

Transformer 诞生时的偏见:随机初始化下的结构归纳偏见及其实际影响

Siquan Li, Yao Tong, Haonan Wang, Tianyang Hu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National University of Singapore(新加坡国立大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 研究揭示了随机初始化的 Transformer 存在结构偏见,通过 SeedPrint 方法区分初始化差异的模型,并解释了注意力机制导致的汇现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04365 2026-02-05 cs.LG 70%

EXaMCaP: Subset Selection with Entropy Gain Maximization for Probing Capability Gains of Large Chart Understanding Training Sets

EXaMCaP: 通过熵增最大化进行子集选择以探测大规模图表理解训练集的探测能力提升

Jiapeng Liu, Liang Li, Bing Li, Peng Fu, Xiyan Gao, Chengyang Fang, Xiaoshuai Hao, Can Ma

机构 * Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所) School of Cyberspace Security, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学网络安全学院) School of Computer and Artificial Intelligence, Jiangxi University of Finance and Economics(江西财经大学计算机与人工智能学院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 EXaMCaP通过熵增最大化选择子集,以高效探测大规模图表理解训练集的能力提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16540 2026-02-04 cs.SD cs.AI eess.AS 70%

Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG

模型是像我们一样听吗?探查音频大语言模型与自然EEG的表征对齐

Haoyun Yang, Xin Xiao, Jiang Zhong, Yu Tian, Dong Xiaohua, Yu Mao, Hao Wu, Kaiwen Wei

机构 * School of Computer Science, Chongqing University(重庆大学计算机科学学院) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(清华大学计算机科学与技术系、人工智能研究院) School of Economics and Business Administration, Chongqing University(重庆大学经济与商业管理学院) School of Artificial Intelligence, Southwest University(西南大学人工智能学院) The First Affiliated Hospital of Chongqing Medical University(重庆医科大学第一附属医院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究通过比较音频大语言模型与EEG信号,揭示了模型在自然聆听中的表征对齐特性,发现排名依赖分裂、时空对齐模式及情感分离现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21875 2026-02-04 cs.CL 70%

LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals

LUMINA:通过上下文-知识信号检测RAG系统中的幻觉

Samuel Yeh, Sharon Li, Tanwi Mallick

机构 * Department of Computer Science, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系) Argonne National Laboratory(阿贡国家实验室)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 LUMINA通过上下文-知识信号检测RAG系统中的幻觉,利用分布距离和token演变测量,实现高准确率和实用性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12387 2026-02-04 cs.LG cond-mat.dis-nn cond-mat.stat-mech math-ph math.MP q-bio.NC stat.ML 70%

Neural Thermodynamics: Entropic Forces in Deep and Universal Representation Learning

神经热力学:深度和通用表征学习中的熵力

Liu Ziyin, Yizhou Xu, Isaac Chuang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出神经热力学理论,揭示深度学习中熵力与对称性打破对表征学习和优化行为的调控作用。

Comments Published at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02018 2026-02-03 cs.AI 70%

Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction

我真的知道吗?学习事实自我验证以减少幻觉

Enes Altinisik, Masoomali Fatehkia, Fatih Deniz, Nadir Durrani, Majd Hawasly, Mohammad Raza, Husrev Taha Sencar

机构 * Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa University (HBKU), Doha, Qatar. \,The authors contributed equally. \,The authors jointly supervised the work

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 VeriFY通过训练时的自我验证框架减少LLM的事实性幻觉,通过一致性推理和阶段级损失屏蔽提升准确性与召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01206 2026-02-03 cs.AI 70%

Addressing Explainability of Generative AI using SMILE (Statistical Model-agnostic Interpretability with Local Explanations)

通过SMILE(统计模型无关可解释性与局部解释)解决生成AI的可解释性问题

Zeinab Dehghani

机构 * School of Computer Science(计算机科学学院) University of Hull(赫尔大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 gSMILE通过受控扰动、Wasserstein距离和加权替代模型,提升生成AI的可解释性,实现细粒度归因和直观热图生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00269 2026-02-03 cs.CV cs.AI 70%

FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering

FaithSCAN: 基于模型的单次幻觉检测用于可靠的视觉问答

Chaodong Tong, Qi Zhang, Chen Li, Lei Jiang, Yanbing Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences (CAS) and the School of Cyber Security, University of CAS(信息工程研究所、中国科学院(CAS)和安全学院、CAS大学)

专题命中 知识编辑与模型理解 :LLM(abstract);language model(abstract);分类 cs.AI

AI总结 FaithSCAN通过利用VLMs的内部信号实现高效准确的视觉问答幻觉检测,克服现有方法的效率与鲁棒性限制。

Comments 21 pages, 13 figures, 8 tables. Submitted to IEEE Transactions on Big Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04753 2026-02-02 cs.CL 70%

EtCon: Edit-then-Consolidate for Reliable Knowledge Editing

EtCon:编辑后整合以实现可靠的知识编辑

Ruilin Li, Yibin Wang, Wenhong Zhu, Chenglin Li, Jinghao Zhang, Chenliang Li, Junchi Yan, Jiaqi Wang

机构 * Wuhan University(武汉大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 EtCon通过编辑后整合范式提升LLMs的知识编辑可靠性与现实应用能力,保留预训练能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07775 2026-02-02 cs.CL 70%

The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs

AI对齐中的意外权衡:在LLMs中平衡幻觉缓解与安全

Omar Mahmoud, Ali Khalil, Buddhika Laknath Semage, Thommen George Karimpanal, Santu Rana

机构 * Applied Artificial Intelligence Initiative(应用人工智能倡议) Deakin University(德肯大学) School of Information Technology(信息科技学院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出了一种方法,通过解耦幻觉和拒绝特征,平衡LLMs中的真实性与安全性,减少因提高真实性而削弱安全对齐的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22595 2026-02-02 cs.AI 70%

Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR

学得更多,用得更少:不确定性一致性引导的查询选择用于RLVR

Hao Yi, Yulan Hu, Xin Li, Sheng Ouyang, Lizhong Ding, Yong Liu

机构 * Renmin University of China(中国人民大学) Gaoling School of Artificial Intelligence(北京人工智能学院) Amap, Alibaba Group(阿里集团阿里的地图部门) School of Computer Science & Technology(计算机科学与技术学院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出基于不确定性一致性的查询选择方法,通过改进RLVR的样本选择策略,减少查询预算,提升推理任务的效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20476 2026-01-29 cs.CL 70%

Can We Improve Educational Diagram Generation with In-Context Examples? Not if a Hallucination Spoils the Bunch

用上下文示例能改进教育图表生成吗?如果幻觉破坏了整体效果就不能

Evanfiya Logacheva, Arto Hellas, Tsvetomila Mihaylova, Juha Sorva, Ava Heinonen, Juho Leinonen

机构 * Aalto University(阿尔托大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过基于RST的上下文示例方法改进教育图表生成,发现高复杂度上下文增加幻觉风险,LLMs难以检测错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17397 2026-01-27 cs.CL 70%

CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing

CLM-Bench: 评估和分析LLMs在知识编辑中的跨语言不对齐

Yucheng Hu, Wei Zhou, Juesi Xiao

机构 * Tianjin University, School of Future Technology(天津大学,未来技术学院) Tianjin University, College of Intelligence and Computing(天津大学,智能计算学院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 CLM-Bench通过构建文化意识的基准,揭示了LLMs在知识编辑中的跨语言不对齐问题,挑战了现有跨语言转移方法的有效性。

Comments EACL MME workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17168 2026-01-27 cs.AI cs.MA 70%

Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability

解释代理系统:超越模型解释到系统级问责

Judy Zhu, Dhari Gandhi, Himanshu Joshi, Ahmad Rezaie Mianroodi, Sedef Akinli Kocak, Dhanesh Ramachandran

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) University of Texas, Austin(德克萨斯大学奥斯汀分校) Dalhousie University(达尔豪斯大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文探讨了代理系统中可解释性技术的必要性,提出需设计专门方法以确保系统在目标形成、环境交互和结果评估等阶段的可追溯性和问责性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05840 2026-01-27 cs.LG 70%

Multimodal Trajectory Representation Learning for Travel Time Estimation

多模态轨迹表示学习用于旅行时间估计

Zhi Liu, Xuyuan Hu, Xiao Han, Zhehao Dai, Zhaolin Deng, Guojiang Shen, Xiangjie Kong

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University of Technology College of Computer Science(浙江工业大学计算机科学学院)

专题命中 知识编辑与模型理解 :language model(abstract);pretraining(abstract);分类 cs.LG

AI总结 本文提出多模态动态轨迹整合框架,通过整合GPS、网格轨迹和道路网络约束,提升旅行时间估计的性能和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15473 2026-01-23 cs.CV cs.LG eess.IV 70%

Emergence and Evolution of Interpretable Concepts in Diffusion Models

扩散模型中可解释概念的涌现与演化

Berk Tinaz, Zalan Fabian, Mahdi Soltanolkotabi

机构 * Dept. of Electrical and Computer Engineering University of Southern California(电气与计算机工程系 美国南加州大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本研究利用SAEs框架揭示扩散模型中可解释概念的涌现与演化,发现早期阶段可控制图像组成,中间阶段确定组成,后期仅能改变细节。

Comments 32 pages, 32 figures, published at the 39th Conference on Neural Information Processing Systems (NeurIPS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13548 2026-01-21 cs.LG 70%

Patterning: The Dual of Interpretability

模式化:可解释性的对偶问题

George Wang, Daniel Murfet

专题命中 知识编辑与模型理解 :language model(abstract);small language model(abstract);分类 cs.LG

AI总结 通过模式化方法,利用敏感性分析确定训练数据以实现特定泛化形式,展示了在语言模型和合成任务中对模型结构的控制与优化。

详情

展开后加载摘要…

URL PDF HTML 收藏