arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7596 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7596 篇

2606.31338 2026-07-01 cs.SD cs.AI eess.AS 新提交 79%

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

超越二元乐器问答:探究音乐音频-语言模型中的乐器接地

Yujun Lee, Joonhyeok Shin, Hyoeun Kim, Kyuhong Shim

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文通过多轴诊断基准测试发现,音乐音频-语言模型在二元乐器问答中的高准确率常掩盖选项位置偏差、易混淆乐器错误和时间响应偏差等问题,表明乐器接地评估需采用多维度基准而非单一准确率。

Comments Workshop on Machine Learning for Audio, ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21815 2026-06-30 cs.CV cs.LG 79%

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

高熵标记作为视觉-语言模型中的多模态失败点

Mengqi He, Xinyu Tian, Xin Shen, Jinhong Ni, Shu Zou, Zhaoyuan Yang, Jing Zhang

机构 * The Australia National University(澳大利亚国立大学) The University of Queensland(昆士兰大学) GE research(GE研究)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 本研究揭示视觉-语言模型中约20%的高熵标记集中了不成比例的对抗性影响,并提出基于熵引导的稀疏攻击方法(EGA),实现高攻击成功率与有害率。

Comments 19 Pages,11 figures,8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28273 2026-06-29 cs.CL 新提交 79%

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

视觉默认,先验覆盖:视觉-语言模型中感知-知识冲突的因果机制

Niclas Lietzow, Danielle Bitterman, Carsten Eickhoff, William Rudman, Michal Golovanevsky

机构 * University of Tübingen(图宾根大学) Harvard University(哈佛大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 通过激活修补和消融实验,发现VLM中视觉默认激活,而先验知识依赖少量因果注意力头(2.5-4.8%),形成不对称因果结构。

Comments 14 pages, 11 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27460 2026-06-29 cs.CL 新提交 79%

Developmental approach reveals the statistical learning of Neural Language Models: Transformers generalize from the most abstract statistical patterns

发展方法揭示神经语言模型的统计学习:Transformer从最抽象的统计模式中泛化

Wang Bojun, Holly Jenkins, Elizabeth Wonnacott

机构 * Department of Education, University of Oxford(牛津大学教育学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 本研究采用发展方法,通过训练生成式Transformer模型并分析其内部表征变化,发现神经语言模型先习得最抽象的全局统计知识,后习得局部统计依赖,并提出新的统计学习与语言认知框架。

Comments 10 pages, 7 figures, oral presentation at Interdisciplinary Advances in Statistical Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24026 2026-06-24 cs.AI 新提交 79%

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

语言模型代理能否成为机械可解释性中有用的电路解释器?

Ayan Antik Khan, Harsh Kohli, Yuekun Yao, Huan Sun, Ziyu Yao

机构 * George Mason University(乔治梅森大学) The Ohio State University(俄亥俄州立大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 研究语言模型代理在机械可解释性中自动解释已定位电路的能力,提出HyVE方法,通过迭代观察、假设生成和因果验证生成组件级解释,实验表明代理有潜力但可靠验证仍是关键障碍。

Comments 23 pages, 4 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20433 2026-06-23 cs.CL 版本更新 79%

Disentangling Geometry, Performance, and Training in Language Models

解耦语言模型中的几何、性能与训练

Atharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian, Swabha Swayamdipta

机构 * University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 系统研究Transformer权重几何(尤其是解嵌入矩阵有效秩)与下游性能的关系,发现有效秩主要反映训练超参数而非性能,不能可靠预测模型表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00918 2026-06-23 cs.AI 版本更新 79%

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

稀疏神经元消融引发大型视觉-语言模型中语言核心的灾难性崩溃

Cen Lu, Yung-Chen Tang, Andrea Cavallaro

机构 * École Polytechnique Fédérale de Lausanne (EPFL)(瑞士洛桑联邦理工学院) Idiap Research Institute(日内瓦研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出渐进式神经元消融方法CAN,发现消融极少量(如LLaVA-1.5-7b中仅4个)位于语言模型下投影层的神经元即可导致LVLM灾难性崩溃,揭示功能依赖语言骨干中的稀疏神经元子集。

Comments Accepted to ICML 2026 Mechanistic Interpretability Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19815 2026-06-19 cs.CL 新提交 79%

Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability

聚类即一切:利用语言模型中的语义聚类预训练Tsetlin Machine以实现可解释性

Jiechao Gao, Rohan Kumar Yadav, Yuangang Li, Yuandong Pan, Jie Wang, Ying Liu, Michael Lepech

机构 * Independent Researcher(独立研究员) University of California, Irvine(加州大学尔湾分校) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 提出一种语义预训练框架,通过K-means或Top2Vec将文本聚类,用聚类-样本对预训练Tsetlin Machine,使其学习可解释的语义关键词,在五个数据集上性能优于传统方法且与BERT竞争。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17417 2026-06-17 cs.SD cs.LG 新提交 79%

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

大型音频语言模型时间理解失败模式的深入分析

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Sarah Wiegreffe, Dinesh Manocha, Ramani Duraiswami

机构 * University of Maryland, College Park(马里兰大学帕克分校)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 本文通过行为与因果机制分析,揭示大型音频语言模型在时间推理中因模态不平衡而失败,并提出注意力重分配方法提升准确率。

Comments Accepted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16044 2026-06-16 cs.LG q-bio.QM 新提交 79%

Circuit Tracing in Autoregressive Protein Language Models

自回归蛋白质语言模型中的电路追踪

Darin Tsui, William Deinzer, Daniel Saeedi, Amirali Aghazadeh

机构 * Stanford University(斯坦福大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 提出ProGenMech框架,通过跨层稀疏编码器忠实恢复ProGen3的生成计算,并零样本发现与蛋白质生成和适应性预测相关的稀疏电路,揭示生物意义基序。

Comments Accepted into the Mechanistic Interpretability Workshop at ICML 2026. 24 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14373 2026-06-15 hep-ex cs.LG hep-ph physics.data-an physics.ins-det 新提交 79%

Machine-learned particle flow as a foundation model for collider physics

机器学习粒子流作为对撞机物理学的基础模型

Farouk Mokhtar, Joosep Pata, Michael Kagan, Javier Duarte

机构 * University of California San Diego(加州大学圣地亚哥分校) National Institute of Chemical Physics and Biophysics(化学物理与生物物理国家研究所) SLAC National Accelerator Laboratory(斯坦福线性加速器中心国家加速器实验室)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 将事件重建视为机器学习问题,利用MLPF模型学习到的潜在表示,在喷注味识别、喷注能量回归和缺失动量回归三项分析任务上显著提升性能,且单线性层即可媲美先进架构,参数减少约35倍。

Comments 15 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06037 2026-06-10 cs.SD cs.CL eess.AS 交叉投稿 79%

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

SpeechJBB:探究大型音频语言模型在代码切换语音下的安全对齐与理解

Virginia Ceccatelli, Yejin Jeon, David Ifeoluwa Adelani

机构 * Mila - Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 提出SpeechJBB数据集,通过代码切换有害音频和伪词插入方法,揭示大型音频语言模型在多语言和口语设置下的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08394 2026-06-09 cs.CL 新提交 79%

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models

当正确决策隐藏内部压力:多模态语言模型中的决策状态探测

Haoran Zhao, Soyeon Caren Han, Eduard Hovy

机构 * The University of Melbourne(墨尔本大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 提出S³E框架,通过正锚定A/B强制选择任务和隐藏状态分析,发现多模态语言模型在正确行为下仍存在语义压力导致的决策状态位移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07861 2026-06-09 cs.CV cs.AI 新提交 79%

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

最后一个可见像素:探究视觉-语言模型中的精细尺度感知

Lujun Li, Lama Sleem, Niccolo Gentile, Yangjie Xu, Yewei Song, Wenbo Wu, Radu State

机构 * University of Luxembourg(卢森堡大学) Foyer S.A. Université Paris-Saclay(巴黎-萨克雷大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出FineSightBench基准,通过4-48像素尺度分离感知与推理任务,发现视觉-语言模型感知在12像素饱和,推理在更大尺度仍受限,揭示精细视觉推理的根本缺陷。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07771 2026-06-09 astro-ph.IM astro-ph.GA cs.AI 新提交 79%

Beyond Point Estimates: Benchmarking Uncertainty Quantification Methods on the AION-1 Astronomical Foundation Model

超越点估计:在AION-1天文基础模型上基准测试不确定性量化方法

Karla Tame-Narvaez, Aleksandra Ćiprijanović, Shubhendu Trivedi

机构 * Scientific Computing Division Fermi National Accelerator Laboratory(费米国家加速器实验室科学计算部) Fermi National Accelerator Laboratory(费米国家加速器实验室) Department of Astronomy and Astrophysics University of Chicago(芝加哥大学天文学与天体物理学系) NSF and Simons SkAI Institute(国家科学基金会与Simons SkAI研究所) Google DeepMind(谷歌DeepMind)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 本文在AION-1基础模型嵌入上比较七种不确定性量化方法,发现共形预测(尤其是LVD框架)在星系属性回归中提供可靠的边际和局部覆盖,优于非共形基线。

Comments 7 pages, 1 table, 1 figure

Journal ref Contribution to Conference on Physics and AI at Stanford University (PAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07025 2026-06-08 cs.CV cs.AI 版本更新 79%

The Geometry of Representational Failures in Vision Language Models

视觉语言模型中表征失败的几何结构

Daniele Savietto, Declan Campbell, André Panisson, Marco Nurisso, Giovanni Petri, Jonathan D. Cohen, Alan Perotti

机构 * Dipartimento di Fisica, Università di Torino(都灵大学物理系) Princeton Neuroscience Institute and AI Lab, Princeton University(普林斯顿大学神经科学研究所和AI实验室) Intesa Sanpaolo AI Research(Intesa Sanpaolo AI研究中心) Dipartimento di Scienze Matematiche, Politecnico di Torino(都灵理工学院数学科学系) Network Science Institute, Northeastern University London, UK(伦敦大学东北方大学网络科学研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 通过分析开源视觉语言模型的概念向量几何重叠,揭示多目标视觉任务中幻觉等错误与认知约束的关联,并提出基于干预的验证方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06273 2026-06-05 cs.IT cs.AI math.IT 79%

Adapting Diffusion Language Models for Lossless Pixel-Level Image Transmission

适应扩散语言模型用于无损像素级图像传输

Tianqi Ren, Rongpeng Li, Xianfu Chen, Yingyu Li, Zhifeng Zhao

机构 * College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) Shenzhen CyberAray Network Technology Co., Ltd(深圳CyberAray网络技术有限公司) School of Mechanical Engineering and Electronic Information, China University of Geosciences(中国地质大学(武汉)机械与电子信息学院) Zhejiang Lab(浙江实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出基于离散扩散模型的分离源信道编码框架DDM-SSCC,通过双向注意力下的同步逆向算术编码实现无损像素级图像传输,并引入Halton引导去噪顺序、掩码率感知余弦调度和轻量温度校准模块提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04180 2026-06-04 cs.LG cs.IT math.IT 79%

KODA: Contrastive Representation Comparison and Alignment for Vision-Language Foundation Models

KODA: 视觉-语言基础模型的对比表示比较与对齐

Youqi Wu, Mohammad Jalali, Farzan Farnia

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 提出KODA框架,通过核优化方法对比分析视觉-语言基础模型的表示差异,并识别弱聚类与强聚类的样本子集,实现表示对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16331 2026-06-04 q-bio.BM cs.AI 79%

Retrieval and competition: how a protein foundation model starts a protein

检索与竞争:蛋白质基础模型如何启动蛋白质

Piotr Jedryszek, Oliver M. Crook

机构 * Department of Biology, University of Oxford, Oxford, UK(牛津大学生物学系) Kavli Institute for Nanoscience Discovery, University of Oxford, Oxford, UK(牛津大学纳科学发现研究所) Department of Chemistry, University of Oxford, Oxford, UK(牛津大学化学系)

专题命中 知识编辑与模型理解 :foundation model(title);language model(abstract);分类 cs.AI

AI总结 通过追踪ESM2-8M预测蛋白质起始甲硫氨酸的计算路径,发现模型依赖位置先验检索而非直接识别,揭示了模型置信度与生物学证据之间的脱节。

Comments updated figure 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13796 2026-06-04 cs.CL cs.CV 79%

The Mechanistic Emergence of Symbol Grounding in Language Models

语言模型中符号接地机制的涌现

Shuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang, Josue Torres-Fonseca, Freda Shi, Joyce Chai

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 通过机械因果分析,发现符号接地在语言模型的中层计算中通过注意力头聚合环境信息实现,并在多模态对话和多种架构中复现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03924 2026-06-03 cs.CL 79%

Knowledge Editing in Masked Diffusion Language Models

掩码扩散语言模型中的知识编辑

Haewon Park, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(首尔国立大学数据科学研究生院)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 研究将定位-编辑方法从自回归模型迁移到掩码扩散模型,发现编辑位置可迁移但多词编辑性能下降,并提出优化中间状态的简单修正方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02374 2026-06-02 cs.AI 79%

Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models

超越像素的空间表示学习:统一栅格数据和向量语义以构建以人为中心的地理空间基础模型

Steffen Knoblauch, Hao Li, Gengchen Mai, Konstantin Klemmer, Song Gao, WenWen Li

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 本文提出统一栅格感知与向量推理的联合空间表示学习范式,旨在解决当前地球观测基础模型仅依赖栅格模态、忽略向量数据中丰富结构化信息的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01899 2026-06-02 eess.SP cs.AI 79%

RA-LWLM: Retrieval-Augmented In-Context Localization with Wireless Foundation Models

RA-LWLM:基于检索增强的上下文无线定位基础模型

Guangjin Pan, Hui Chen, Hei Victor Cheng, Henk Wymeersch

机构 * Department of Electrical Engineering, Chalmers University of Technology(查尔姆斯理工大学电子工程系) Department of Electrical and Computer Engineering, Aarhus University(阿鲁斯大学电子与计算机工程系)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 提出RA-LWLM框架,通过将场景特定信息外化到指纹数据库,实现无需训练的跨场景无线定位,利用冻结的无线基础模型编码器、检索模块和基于Transformer的上下文学习模块预测用户位置。

Comments 13 pages, 9 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01427 2026-06-02 stat.ML cs.LG 79%

On the Uncertainty Quantification Ability of Tabular Foundation Models

关于表格基础模型的不确定性量化能力

Tyler R. Johnson, Kian Ben-Jacob, Nima Negarandeh, Oriol Vendrell-Gallart, Ramin Bostanabad

机构 * Department of Mechanical and Aerospace Engineering, University of California, Irvine(加州大学欧文分校机械与航空航天工程系) Department of Civil and Environmental Engineering, University of California, Irvine(加州大学欧文分校土木与环境工程系)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 通过对比TabPFN与高斯过程在回归任务上的实证研究,揭示了显式先验与学习先验之间的权衡:TabPFN在复杂高维问题中表现优异,而高斯过程在数据稀缺时提供更优的预测精度和不确定性量化。

Comments 12 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30865 2026-06-01 cs.LG 79%

GlucoFM: A Dual-Stream Foundation Model for Continuous Glucose Monitoring

GlucoFM: 一种用于连续血糖监测的双流基础模型

Zechen Li, Keerthana Natarajan, Weizhi Zhang, Menglian Zhou, Simon A. Lee, Yuwei Zhang, Maxwell A. Xu, Zeinab Esmaeilpour, Flora D. Salim, Mark Malhotra, Lindsey Sunden, Shwetak Patel, Yuzhe Yang, Ahmed A. Metwally

机构 * Google Research(谷歌研究) University of New South Wales(新南威尔士大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 提出GlucoFM,一种轻量级CGM基础模型,通过将血糖动态分解为慢生理状态和瞬态事件流,在7个临床预测任务上平均PR-AUC比最佳CGM专用模型提高4.1点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29881 2026-05-29 cs.CV cs.AI 79%

Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering

通过屏障调控自适应闭式引导缓解视觉语言模型中的幻觉

Soumyadeep Jana, Pulkit Mittal, Sanasam Ranbir Singh

机构 * Indian Institute of Technology Guwahati(印度理工学院果阿班加)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出BRACS框架,通过监测视觉注意力并仅在接地退化时进行闭式修正,无需训练即可有效减少LVLM中的物体幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23853 2026-05-29 cs.AI cs.MA 79%

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

SCoOP: 多视觉-语言模型系统中用于不确定性量化的语义一致意见池化

Chung-En Johnny Yu, Brian Jalaian, Nathaniel D. Bastian

机构 * University of West Florida(西佛罗里达大学) United States Military Academy(美国军事学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出SCoOP框架,通过不确定性加权的线性意见池化聚合多个视觉-语言模型的输出,实现无训练的不确定性量化,有效检测幻觉并支持高不确定性样本的弃权。

Comments Accepted to ICLR 2026 Workshop on Agentic AI in the Wild: From Hallucinations to Reliable Autonomy

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28554 2026-05-28 cs.LG 79%

High Performance, Low Reliability: Uncertainty Benchmarking for Tabular Foundation Models

高性能,低可靠性:表格基础模型的不确定性基准测试

José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel, Bich-Liên Doan

机构 * CentraleSupélec(中央理工大学) ENS Paris-Saclay(巴黎-萨克雷大学) Université Paris-Saclay(巴黎-萨克雷大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 通过TALENT基准测试,发现表格基础模型虽在预测性能上优于梯度提升决策树,但在不确定性校准上表现更差,存在性能-不确定性权衡。

Comments 6 pages, 2 figures, 2 tables. Accepted at ESANN 2026 (European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning), 22-24 April 2026, Bruges (Belgium)

Journal ref ESANN 2026 proceedings, European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, Bruges (Belgium) and online event, 22-24 April 2026, pp. 115-120, i6doc.com publ., ISBN 9782875870964

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00913 2026-05-28 cs.CV cs.CL 79%

Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

跨描绘装配指令对齐的视觉-语言模型基准测试与机制分析

Zhuchenyang Liu, Yao Zhang, Yu Xiao

机构 * Aalto University(阿alto大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 构建IKEA-Bench基准,评估19个视觉-语言模型在装配图与视频帧对齐任务上的表现,发现视觉编码是提升跨描绘鲁棒性的关键瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07961 2026-05-26 cs.AI 79%

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

探究语言模型的偏好:整合AI福祉的言语与行为测试

Valen Tagliabue, Leonard Dung

机构 * Future Impact Group (FIG)(未来影响集团) Ruhr-University Bochum(波鸿鲁尔大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本研究通过言语报告和行为实验(虚拟环境导航与话题选择)测量语言模型的偏好,发现偏好满足可作为AI福祉的实证代理,但测量一致性因模型和条件而异。

Comments Forthcoming in Philosophy and the Mind Sciences (PhiMiSci)

详情

展开后加载摘要…

URL PDF HTML 收藏