arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-09-01 至 2026-09-01 共收录 54 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 54 篇

2606.01914 2026-09-01 cs.CL cs.CV 版本更新 92%

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

多模态大语言模型空间推理中空间词汇偏差的机制诊断

Chuang Ma, Qianying Liu, Tomoyuki Obuchi, Fei Cheng, Wang Yang, Sudong Cai, Shuyuan Zheng, Akiko Aizawa, Sadao Kurohashi

机构 * Kyoto University(京都大学) NII LLMC(日本国立信息与通信技术研究所语言模型中心) RIKEN AIP(日本理化学研究所先进理工研究所) Case Western Reserve University(凯斯西储大学) The Hong Kong Polytechnic University(香港理工大学) The University of Osaka(大阪大学) University of Tokyo(东京大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文发现多模态大语言模型存在空间词汇偏差,即添加空间关系词会吸引模型选择该选项,并通过机制可解释性工具揭示偏差主要源于语言侧而非视觉侧,最后提出轻量级LLM-only DPO更新可有效缓解偏差。

Comments 27 pages. Accepted to EMNLP 2026 (Main Conference); camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10107 2026-09-01 cs.HC 版本更新 91%

The Double-Edged Sword of Open-Ended Interaction: How LLM-Driven NPCs Affect Players' Cognitive Load and Gaming Experience

开放式交互的双刃剑:LLM驱动的NPC如何影响玩家的认知负荷和游戏体验

Ting-Chen Hsu, Wenran Chen, Jiangxu Lin, Fei Qin, Zheyuan Zhang

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 研究探讨LLM驱动NPC对玩家认知负荷和游戏体验的影响,揭示心理机制、任务场景差异及个体特质作用,发现NPC显著增加认知负荷但未提升整体体验,且效果随任务场景变化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02182 2026-09-01 cs.LG cs.CL 版本更新 91%

Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation

贝叶斯稀疏低秩自适应用于大语言模型不确定性估计

Jijie Zhang, Zhe Ren, Quan Zhang, Dandan Guo

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Michigan State University(密歇根州立大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.CL、cs.LG

AI总结 提出DALorRA,一种变分贝叶斯稀疏框架,通过随机掩码低秩适应中的秩维度实现模型容量正则化和校准,在不牺牲推理精度下提升LLM校准性能。

Comments To appear in EMNLP 2026. 17 pages, 7 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18123 2026-09-01 cs.AI cs.LG 版本更新 91%

SPADE: A Large Language Model Framework for Soil Moisture Pattern Recognition and Anomaly Detection in Precision Agriculture

SPADE:面向精准农业土壤湿度模式识别与异常检测的大语言模型框架

Yeonju Lee, Rui Qi Chen, Joseph Oboamah, Po Nien Su, Wei-zhen Liang, Yeyin Shi, Lu Gan, Yongsheng Chen, Xin Qiao, Jing Li

机构 * organization= H. Milton Stewart School of Industrial Systems Engineering, Georgia Institute of Technology , city= Atlanta , state= GA , country= USA organization= Panhandle Research Extension Center, University of Nebraska-Lincoln , city= Scottsbluff , state= NE , country= USA organization= Department of Computer Science Engineering, University of Nebraska-Lincoln , city= Lincoln , state= NE , country= USA organization= Department of Biological Systems Engineering, University of Nebraska-Lincoln , city= Lincoln , state= NE , country= USA organization= Institute for Robotics Intelligent Machines, Georgia Institute of Technology , city= Atlanta , state= GA , country= USA organization= School of Civil \& Environmental Engineering, Georgia Institute of Technology, Georgia Institute of Technology , city= Atlanta , state= GA , country= USA

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(title);language model(title);分类 cs.AI、cs.LG

AI总结 本研究提出首个基于LLM的土壤湿度时间序列分析框架SPADE,用GPT-4.1结合领域提示零样本识别湿润事件与异常,在真实多作物数据上优于无训练基线,可生成结构化报告辅助土壤湿度解读。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30776 2026-09-01 eess.AS 新提交 89%

Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR

用于基于大语言模型的自动语音识别幻觉缓解的似然约束声学重排序(无需训练)

Jiasheng Kuang, Linru Zheng, Hongjin Song, Zhaoqi Cui, Song Li

专题命中 知识编辑与模型理解 :LLM(title,summary_cn);large language model(abstract);language model(abstract)

AI总结 本研究针对基于LLM的ASR系统的幻觉问题,提出无需训练的LCAR方法,通过似然约束与声学重排序缓解幻觉,在δ=0.60时消除38.8%-57.1%的幻觉,同时维持词/字符错误率。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03985 2026-09-01 cs.CR cs.AI 版本更新 89%

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

NeuroBreak:揭示大语言模型的内部越狱机制

Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan, Tianyu Du, Shouling Ji, Dazhen Deng, Yingcai Wu

机构 * Zhejiang University(浙江大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 NeuroBreak是一个自上而下的大语言模型越狱分析系统,可分析神经元级安全机制、关键神经元,经评估验证了有效性,为开发下一代防御策略提供机制见解。

Comments 11 pages, 10 figures. Accepted to IEEE VIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28740 2026-09-01 cs.CL cs.AI 版本更新 88%

Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text

反向探测:临床文本中大语言模型的监督式词级不确定性量化

Bushi Xiao, Sarvesh Soni, Daisy Zhe Wang

机构 * University of Florida(佛罗里达大学) U.S. National Library of Medicine(美国国家医学图书馆)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 提出反向探测框架,利用预标注摘要从模型内部激活中提取词级不确定性信号,在临床文本中实现高效、可解释的不确定性量化。

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29808 2026-09-01 cs.CR cs.PL cs.SE 新提交 86%

POLYFLOW: A Neuro-Symbolic Framework for Static Cross-Language Information Flow Analysis

POLYFLOW:用于静态跨语言信息流分析的神经符号框架

Haoran Yang, Zhixuan Zhong, Jiawei Guo, Haipeng Cai

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract)

AI总结 PolyFlow是结合LLM与静态分析的跨语言信息流分析神经符号框架,在多语言系统实验中表现优于基线,可发现未知跨语言漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29974 2026-09-01 cs.CV cs.CL 新提交 85%

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

SpanCalib-VLM:视觉语言模型中经过校准的幻觉跨度检测

Amanuel Gizachew Abebe, Yasmin Moslem

机构 * Shaggar Institute of Technology(沙加尔理工学院) Trinity College Dublin(都柏林圣三一学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);SFT(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出SpanCalib-VLM混合双系统,结合多模态序列标注器与微调生成式VLM,经联合校准融合策略优化,在SHROOM-Visions任务上实现幻觉跨度的高效校准检测,公开了模型权重与代码。

Comments Shroom-Visions

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07524 2026-09-01 cs.CL cs.AI 版本更新 85%

ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding

ABLE:基于归因的大模型嵌入表示与映射

Zirui Wang, Yusen Hou, Shaofeng Liang, Bowen Tian, Yanlin Zhang, Wenshuo Chen, Yutao Yue

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Deep Interdisciplinary Intelligence Lab (DI2 Lab)(深度跨学科智能实验室(DI2 Lab))

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出ABLE框架,利用梯度特征归因和分词器无关的词级对齐构建模型嵌入,实现异构LLM的高效比较,在关系预测、模型路由和基准分数预测上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28930 2026-09-01 cs.CL cs.AI cs.LG 新提交 83%

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

幻觉信号是均值漂移:为何简单探测方法就足够

Jungseob Lee, Jaehyung Seo, Heuiseok Lim

机构 * Korea University(高丽大学) Konkuk University(建国大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现LLM幻觉检测的隐藏状态信号以单一均值漂移分量为主,简单的L2正则化逻辑回归等方法即可达到良好性能,LayerMix可聚合信号实现接近先知层的检测效果。

Comments 19 pages, 7 figures, 20 tables. Accepted to EMNLP 2026 (Main Conference). Code: this https URL (https://github.com/js-lee-AI/LayerMix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30751 2026-09-01 cs.AI cs.CV 新提交 83%

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

自回归马赛克:探测仅文本语言模型的二维空间推理能力

Ashwin Nedungadi, Stefan Oehmcke, Stefan Lüdtke

机构 * Institute for Visual & Analytic Computing (VAC), University of Rostock(罗斯托克大学视觉与分析计算研究所(VAC))

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 该研究引入AM-Bench基准,发现仅文本LLMs的二维空间表现取决于模型和输出介质,无法仅用代码生成能力解释,其开放式布局表现存在显著差异。

Comments WACV 2027 Submission Pre-Print

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29431 2026-09-01 cs.CL 新提交 83%

Evaluating the Semantic Specificity of Representation Steering in Language Models

评估语言模型中表征引导的语义特异性

Zhangdie Yuan, Andreas Vlachos

机构 * University of Cambridge(剑桥大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 本研究提出CRT诊断框架,发现LRS仅注入全局标签偏差而非修复推理回路,为区分真正推理修复与表面标签覆盖提供了严谨方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22130 2026-09-01 cs.MA cs.CL 版本更新 83%

PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems

PropUQ-MAS:面向大语言模型多智能体系统的传播感知不确定性量化

Yaokun Liu, Yifan Liu, Daniel Yue Zhang, Ruichen Yao, Zelin Li, Dong Wang

机构 * Scale AI University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.CL

AI总结 针对现有不确定性量化方法无法捕捉大语言模型多智能体系统中不确定性传播的问题,提出PropUQ-MAS框架,实验显示其可显著提升该系统的不确定性量化性能。

Comments Accepted to EMNLP 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01605 2026-09-01 cs.LG stat.ML 版本更新 83%

Universal Redundancies in Time Series Foundation Models

时间序列基础模型中的通用冗余性

Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William Gilpin

机构 * ECE Department, UT Austin(UT奥斯汀电子工程系) Department of Physics, UT Austin(UT奥斯汀物理系) Oden Institute, UT Austin(UT奥斯汀奥登研究所)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);pretraining(abstract);分类 cs.LG

AI总结 本研究揭示了时间序列基础模型中普遍存在的冗余特性,并通过理论框架和消融实验揭示了模型退化现象的根源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29530 2026-09-01 cs.CL cs.AI 新提交 82%

The Emergent Symbolic Structure of Artificial Neural Networks

人工神经网络的涌现符号结构

R. Thomas McCoy, Paul Soulos, Tal Linzen, Paul Smolensky

机构 * Yale University(耶鲁大学) Johns Hopkins University(约翰斯·霍普金斯大学) New York University(纽约大学) Microsoft Research(微软研究院)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究发现神经网络内部表征隐含符号结构,可通过符号结构近似其向量表征,还能精准干预修改大型语言模型行为,为调和智能的符号概念与现代AI的向量本质提供了新途径。

Comments 30 pages, plus 29 pages of references and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14323 2026-09-01 cs.CL 版本更新 81%

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever

超越语义相似度:通过行为对齐检索减少不必要的API调用

Yixin Chen, Ying Xiong, Shangyu Wu, Yufei Cui, Xue Liu, Nan Guan, Chun Jason Xue

机构 * City University of Hong Kong(香港城市大学) Mohamed Bin Zayed University of Artificial Intelligence(马尔代夫穆罕默德·本·扎耶德人工智能大学) McGill University(麦吉尔大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 该研究针对工具增强型LLM的不必要API调用问题,提出行为对齐检索(BAR)方法,通过训练感知行为的相似度排序演示示例,在多类骨干网络和基准中提升调用可靠性、减少API调用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29924 2026-09-01 cs.CV cs.AI cs.LG 新提交 81%

Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization

基于隐式特征稳定化的大视觉语言模型幻觉缓解方法

Aditi Sarker, Rafi Ibn Sultan, Hui Zhu, Dongxiao Zhu, Prashant Khanduri

机构 * Wayne State University(韦恩州立大学) Institute for AI and Data Science (AIDaS), Wayne State University(韦恩州立大学人工智能与数据科学研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 该研究针对大视觉语言模型的幻觉问题,提出隐式特征稳定化框架INFUSE,通过微调内置扰动不变性,在多基准模型上大幅降低幻觉率且无推理开销。

Comments 28 Pages, 12 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29034 2026-09-01 cs.CL cs.AI 新提交 81%

A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

语言模型表示的统一视角:从填充符-角色结构到机制可解释性

Zhang Enyan, R. Thomas McCoy

机构 * Yale University(耶鲁大学) Wu Tsai Institute(吴蔡研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究以张量积表示(TPRs)为统一假设,证明其可统一加法类比等多种语言模型可解释性方法,构建的变体与标准变体性能相当,为神经网络本质的统一阐释奠定基础。

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19926 2026-09-01 cs.CL cs.AI 版本更新 81%

The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models

Transformer的语法:语言模型中句法知识可解释性研究的系统综述

Nora Graichen, Iria de-Dios-Flores, Gemma Boleda

机构 * Universitat Pompeu Fabra(巴塞罗那庞培乌法布拉大学) ICREA(加泰罗尼亚国家研究委员会)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 通过对337篇文章的系统综述,评估基于Transformer的语言模型(TLM)的句法能力,发现TLM编码了非平凡的句法知识,但句法-语义接口现象表现较弱,且研究集中在英语和BERT类模型上。

Comments Published as a main conference paper at EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28924 2026-09-01 cs.CL 新提交 79%

Causal Interventions Reveal Typologically Organized Syntactic Mechanisms in Multilingual Language Models

因果干预揭示多语言语言模型中类型学组织的句法机制

Sasha Boguraev, Toshiki Nakai, Kyle Mahowald, Julius Steuer

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Saarland University(萨尔大学) Leipzig University(莱比锡大学) Heidelberg Institute for Theoretical Studies(海德堡理论研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 本研究利用机制可解释性技术,在四种多语言语言模型的三种句法结构上证实跨语言机制迁移,且迁移程度与语言类型学相似度正相关,为语言学理论提供新假说。

Comments 20 Pages, 7 Figures, 11 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06037 2026-09-01 cs.SD cs.CL eess.AS 版本更新 79%

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

SpeechJBB:探究大型音频语言模型在代码切换语音下的安全对齐与理解

Virginia Ceccatelli, Yejin Jeon, David Ifeoluwa Adelani

机构 * Mila - Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 提出SpeechJBB数据集,通过代码切换有害音频和伪词插入方法,揭示大型音频语言模型在多语言和口语设置下的安全漏洞。

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29936 2026-09-01 cs.CL cs.LG 新提交 79%

When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs

当安全“说”一种语言:大语言模型中安全-语言身份纠缠的机制分析

Apoorva Upadhyaya, Sandipan Sikdar

机构 * L3S Research Center(L3S研究中心) Leibniz Universität Hannover(汉诺威莱布尼茨大学)

专题命中 知识编辑与模型理解 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本研究通过稀疏自编码器特征分析,揭示大语言模型中安全与语言身份的纠缠机制,发现安全相关特征具架构依赖性,为多语言安全干预提供机制解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25100 2026-09-01 cs.AI cs.LG 版本更新 79%

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

面向可靠、可泛化且特定的上下文内知识编辑:多目标强化学习方法

Xuzhong Wang, Maiqi Jiang, Tejal Nair, Girija Bhusal, Yanfu Zhang, Haipeng Chen

机构 * College of William and Mary(威廉玛丽学院) Tribhuvan University(特里布文大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 针对现有上下文内知识编辑方法难以平衡可靠、泛化、特定目标的问题,提出多目标强化学习算法MO-IKE,在Llama-3.2上多项指标较以往方法显著提升。

Comments Our work proposes a multi-objective reinforcement learning algorithm that optimizes prompt construction for reliable, generalizable, and specific in-context knowledge-editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27068 2026-09-01 cs.CL cs.AI cs.MA 版本更新 79%

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

QUACK: 多模态社交推理智能体中的沟通知识质疑、理解与审计

Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) University of Cambridge(剑桥大学) MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学) University of Toronto(多伦多大学) Salesforce

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出QUACK框架,通过游戏结果、行为轨迹和话语一致性三级评估,自动审计多模态社交推理智能体语言与感知行为的一致性,发现最强智能体仍有15.1%的空间幻觉和过半无据指控。

Comments Accepted by EMNLP 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14088 2026-09-01 cs.CV 版本更新 78%

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

VideoRAE:通过表示自动编码器驯服用于生成建模的视频基础模型

Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Huazhong University of Science and Technology(华中科技大学) Shenzhen Loop Area Institute(深圳河套学院) University of Science and Technology of China(中国科学技术大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 研究视频生成模型潜在空间问题,提出VideoRAE,利用冻结视频基础编码器特征经1D自注意力投影仪压缩,支持多种潜在空间,通过多码本高维量化等实现强大重建,收敛快,验证了冻结VFM表示的有效性。

Comments Home page: this https URL (https://zhxie0117.github.io/VideoRAE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30345 2026-09-01 cs.AI 新提交 77%

Answer Probing-Guided Search for Diverse Solution Exploration of LLMs

基于答案探测引导的大语言模型多样化解决方案探索搜索

Yi Fang, Que Shen, Chengpeng Li, Boyi Deng, Wei Shi, Wenjie Wang, Fuli Feng, Fengli Xu, Dayiheng Liu

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Alibaba Group(阿里巴巴集团) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学)

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对LLMs推理易收敛于单一解的问题,提出Answer Probing引导的树搜索APTS,经实验证实可提升多推理任务的解决方案多样性,具备有效性与鲁棒性。

Comments Accepted to the EMNLP 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30646 2026-09-01 cs.CL cs.AI cs.LG 新提交 75%

BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs

BiG-SURE - 用于大语言模型语义不确定性与可靠性估计的二分图

Debarpan Bhattacharya, Malay Phadke, Sriram Ganapathy

机构 * Indian Institute of Science(印度科学学院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 BiG-SURE是一种基于跨温度语义一致性的二分图方法,用于黑盒LLMs/VLMs的语义不确定性与可靠性估计,在多类问答任务上提升了弃权AUROC性能。

Comments 22 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11327 2026-09-01 cs.LG cs.AI 版本更新 73%

PRISM Edit: One Vector for All Temporal Answers

PRISM Edit:适用于所有时间答案的单一向量

Chen Huang, Qi Zheng, Ruiqin Zheng, Long Zeng, Yuantong Xu

机构 * Tsinghua University(清华大学) ByteDance(字节跳动)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究针对大语言模型时间事实更新问题,基于因果追踪发现其内部计算支持新旧答案区分,进而引入PRISM Edit,通过优化单一多义词表示及利用固有调制路径,在新基准上评估,相比基线提升了时间一致性等指标且速度更快。

Comments Chen Huang and Qi Zheng contributed equally. Corresponding authors: Long Zeng, Yuantong Xu

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12868 2026-09-01 cs.CL cs.LG 版本更新 73%

Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues

种族、族裔及其在大语言模型中的偏见影响

Shiyue Hu, Ruizhe Li, Yanjun Gao

机构 * University of Colorado Anschutz(科罗拉多大学安施茨分校) University of Colorado Boulder(科罗拉多大学博尔德分校) University of Aberdeen(阿伯丁大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 研究探讨了种族和族裔在大语言模型中的表示和偏见机制,通过实验发现人口信息在模型内部的分布差异,并提出干预方法以减少偏见影响。

Comments EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏