arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7608 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7608 篇

2601.12690 2026-01-21 cs.HC 50%

"Are we writing an advice column for Spock here?" Understanding Stereotypes in AI Advice for Autistic Users

我们在这里是在为斯波克写建议专栏吗?理解为自闭症用户提供的AI建议中的刻板印象

Caleb Wohn, Buse Çarık, Xiaohan Ding, Sang Won Lee, Young-Ho Kim, Eugenia H. Rho

专题命中 知识编辑与模型理解 :LLM(abstract)

AI总结 研究探讨了自闭症用户在请求AI建议时披露身份对建议生成的影响,发现LLM在披露身份时更倾向于避免刻板印象情境,引发对个性化建议有效性的讨论。

Comments Accepted to CHI '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12481 2026-01-21 cs.CV cs.GR 50%

NeuralFur: Animal Fur Reconstruction From Multi-View Images

NeuralFur: 从多视角图像中重建动物毛发

Vanessa Sklyarova, Berna Kabadayi, Anastasios Yiannakidis, Giorgio Becherini, Michael J. Black, Justus Thies

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ETH Zürich(苏黎世联邦理工学院) Technical University of Darmstadt(德累斯顿技术大学) University of Tübingen(图宾根大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 NeuralFur通过视觉语言模型指导多视角图像重建,实现不同动物的高保真3D毛发建模。

Comments For additional results and code, please refer to https://neuralfur.is.tue.mpg.de

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12479 2026-01-21 cs.RO 50%

Language-Based Swarm Perception: Decentralized Person Re-Identification via Natural Language Descriptions

基于语言的群体感知:通过自然语言描述实现去中心化的人员重识别

Miquel Kegeleirs, Lorenzo Garattoni, Gianpiero Francesca, Mauro Birattari

机构 * IRIDIA, Université libre de Bruxelles(IRIDIA,布鲁塞尔自由大学) Toyota Motor Europe(丰田欧洲 motors)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出一种基于自然语言的去中心化人员重识别方法,通过视觉-语言模型生成文本描述,实现群体协作识别与可解释性提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11243 2026-01-19 cs.CV 50%

Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification

图像-文本知识建模用于无监督多场景人物重识别

Zhiqi Pang, Lingling Zhao, Yang Liu, Chunyu Wang, Gaurav Sharma

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出图像-文本知识建模用于无监督多场景人物重识别,通过三阶段框架提升跨场景识别性能。

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08036 2026-01-14 cs.SE 50%

Automating API Documentation from Crowdsourced Knowledge

从众包知识自动化生成API文档

Bonan Kou, Zijie Zhou, Muhao Chen, Tianyi Zhang

专题命中 知识编辑与模型理解 :LLM(abstract)

AI总结 AutoDoc通过从Stack Overflow提取API知识,生成更准确、简洁且全面的API文档,有效解决官方文档过时和不完整的问题。

Comments 13 pages, 2 figures, Accepted to ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05939 2026-01-12 cs.CV 50%

Context-Aware Decoding for Faithful Vision-Language Generation

面向上下文的解码方法用于忠实的视觉-语言生成

Mehrdad Fazli, Bowen Wei, Ziwei Zhu

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出了一种无需训练的上下文嵌入注入方法,通过利用上下文嵌入信号在解码过程中保持视觉一致性,有效减少视觉-语言模型中的幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05721 2026-01-12 cs.SE 50%

From Issues to Insights: RAG-based Explanation Generation from Software Engineering Artifacts

从问题到洞察:基于RAG的软件工程制品解释生成

Daniel Pöttgen, Mersedeh Sadeghi, Max Unterbusch, Andreas Vogelsang

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文首次应用RAG方法从问题跟踪数据生成解释,实现90%的人工解释一致性,展示了结构化问题数据在提升软件系统可解释性方面的潜力。

Comments Accepted at NLBSE 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03070 2026-01-07 cs.RO 50%

HEXAR: a Hierarchical Explainability Architecture for Robots

HEXAR:机器人中的分层可解释性架构

Tamlin Love, Ferran Gebellí, Pradip Pramanick, Antonio Andriella, Guillem Alenyà, Anais Garrell, Raquel Ros, Silvia Rossi

机构 * Institut de Robòtica i Informàtica Industrial (CSIC-UPC)(机器人与信息工业研究所(CSIC-UPC)) PAL Robotics(PAL机器人技术公司) University of Naples Federico II(那不勒斯费德里科二世大学) Artificial Intelligence Research Institute (IIIA-CSIC)(人工智能研究所(IIIA-CSIC))

专题命中 知识编辑与模型理解 :LLM(abstract)

AI总结 HEXAR通过分层可解释性架构提升机器人系统的透明度和可解释性,有效提高根因识别和运行效率。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08002 2026-01-07 cs.CV 50%

Aligning Text, Images, and 3D Structure Token-by-Token

对齐文本、图像和3D结构逐个标记

Aadarsh Sahoo, Vansh Tibrewal, Georgia Gkioxari

机构 * California Institute of Technology(加州理工学院)

专题命中 知识编辑与模型理解 :LLM(abstract)

AI总结 本文提出统一的LLM框架,实现文本、图像和3D结构的对齐,通过结构化3D场景模态提升3D场景理解与重建能力。

Comments Project webpage: https://glab-caltech.github.io/kyvo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20032 2025-12-30 cs.CV 50%

VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance

VALLR-Pin:基于拼音引导的中文视觉语音识别中的不确定性因子化方法

Chang Sun, Dongliang Xie, Wanpeng Xie, Bo Qin, Hong Yang

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学) First Research Institute of the Ministry of Public Security of the People’s Republic of China(中华人民共和国公安部第一研究所)

专题命中 知识编辑与模型理解 :LLM(abstract)

AI总结 VALLR-Pin通过引入拼音作为中间表示,结合轻量LLM细化模块,提升了中文视觉语音识别的转录准确性,尤其在多说话人条件下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21331 2025-12-29 cs.CV 50%

TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning

TICON:用于病理学表示学习的滑片级上下文化器

Varun Belagali, Saarthak Kapse, Pierre Marza, Srijan Das, Zilinghan Li, Sofiène Boutaj, Pushpak Pati, Srikar Yellapragada, Tarak Nath Nandi, Ravi K Madduri, Joel Saltz, Prateek Prasanna, Stergios Christodoulidis, Maria Vakalopoulou, Dimitris Samaras

机构 * Stony Brook University(石溪大学) MICS, CentraleSupélec, Université Paris-Saclay(MICS、CentraleSupélec、巴黎-萨克雷大学) UNC Charlotte(北卡罗来纳大学夏洛特分校) Argonne National Laboratory(阿贡国家实验室) University of Chicago(芝加哥大学) Archimedes/Athena RC Independent Researcher(独立研究者)

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 TICON通过统一的Transformer模型提升病理学滑片和全滑片表示学习的性能,实现多个基准的新状态-of-the-art结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14330 2025-12-29 cs.IR 50%

Ensembling Multiple Hallucination Detectors Trained on VLLM Internal Representations

集成多个基于VLLM内部表示的幻觉检测器

Yuto Nakamizo, Ryuhei Miyazato, Hikaru Tanabe, Ryuta Yamakura, Kiori Hatanaka

专题命中 知识编辑与模型理解 :LLM(abstract)

AI总结 本文提出通过集成多个基于VLLM内部表示的幻觉检测模型,以减少幻觉并提高VQA任务的准确性。

Comments 5th place solution at Meta KDD Cup 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19347 2025-12-25 physics.optics cs.GR eess.IV 50%

High contrast holography through dual modulation

高对比度全息成像通过双调制

Leyla Kabuli, Oliver Cossairt, Florian Schiffers, Nathan Matsuda, Grace Kuo

专题命中 知识编辑与模型理解 :SLM(abstract)

AI总结 本文提出通过双调制技术提升全息显示对比度,实验显示对比度显著提高,为高对比度全息显示提供新设计思路。

Comments 24 pages, 17 figures

Journal ref Nature Scientific Reports 15, 17615 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16073 2025-12-23 cs.CV cs.RO 50%

OW-Rep: Open World Object Detection with Instance Representation Learning

OW-Rep: 开放世界目标检测中的实例表示学习

Sunoh Lee, Minsik Jeon, Jihong Min, Junwon Seo

机构 * KAIST(韩国科学技术院) Carnegie Mellon University(卡内基梅隆大学) Agency for Defense Development(国防发展局)

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 OW-Rep通过引入实例表示学习,提升开放世界目标检测中未知目标的检测精度和语义嵌入质量,增强下游任务表现。

Comments Accepted to WACV 2026. Our project website can be found at https://sunohlee.github.io/OW-Rep/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15611 2025-12-18 cs.CV 50%

If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions

如果你能描述它,他们就能看到它:从文本描述中跨模态学习视觉概念

Carlo Alberto Barbano, Luca Molinaro, Massimiliano Ciranni, Emanuele Aiello, Vito Paolo Pastore, Marco Grangetto

机构 * University of Turin(都灵大学) University of Genoa(热那亚大学) Politecnico di Torino(都灵理工大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出通过文本描述跨模态学习视觉概念的方法,利用知识迁移技术提升视觉-语言模型的零样本性能。

Comments 27 pages. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13434 2025-12-16 eess.IV cs.CV 50%

Self-Supervised Ultrasound Representation Learning for Renal Anomaly Prediction in Prenatal Imaging

自监督超声表示学习用于产前影像中肾异常预测

Youssef Megahed, Inok Lee, Robin Ducharme, Kevin Dick, Adrian D. C. Chan, Steven Hawken, Mark C. Walker

机构 * organization= Department of Systems Computer Engineering, Carleton University , city= Ottawa , state= Ontario , country= Canada organization= Department of Methodological Implementation Research, Ottawa Hospital Research Institute , city= Ottawa , state= Ontario , country= Canada organization= Department of Acute Care Research, Ottawa Hospital Research Institute , city= Ottawa , state= Ontario , country= Canada organization= Children's Hospital of Eastern Ontario Research Institute , city= Ottawa , state= Ontario , country= Canada organization= Better Outcomes Registry \& Network Ontario, Children’s Hospital of Eastern , city= Ottawa , state= Ontario , country= Canada organization= Department of Obstetrics Gynecology, University of Ottawa , city= Ottawa , state= Ontario , country= Canada organization= School of Epidemiology Public Health, University of Ottawa , city= Ottawa , state= Ontario , country= Canada organization= Department of Obstetrics, Gynecology \& Newborn Care, The Ottawa Hospital , city= Ottawa , state= Ontario , country= Canada Global Health Office, University of Ottawa , city= Ottawa , state= Ontario , country= Canada organization= Department of Clinical Science Translational Medicine, University of Ottawa , city= Ottawa , state= Ontario , country= Canada

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 本文提出了一种自监督超声基础模型,用于产前影像中肾异常的自动分类,通过实验验证该模型在二分类和多类分类任务中均优于传统方法。

Comments 14 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26005 2025-12-15 nlin.CD 50%

Regime identification and control of extremes in the non-autonomous Lorenz model with chaos and intransitivity

非自治洛伦兹模型中混沌与不稳定性下的极端现象识别与控制

Moyan Liu, Qin Huang, Upmanu Lall

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 本文提出了一种基于非均匀隐马尔可夫模型和局部李雅普诺夫指数的自适应混沌控制策略,用于控制季节性驱动和噪声扰动的洛伦兹84模型中的极端现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04552 2025-12-15 eess.SP 50%

DAS-MAE: A self-supervised pre-training framework for universal and high-performance representation learning of distributed fiber-optic acoustic sensing

DAS-MAE:一种用于分布式光纤声学传感通用性和高性能表示学习的自监督预训练框架

Junyi Duan, Jiageng Chen, Zuyuan He

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 DAS-MAE通过自监督预训练框架实现对分布式光纤声学传感信号的高效表示学习,提升少样本分类性能和实际应用中的识别精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08700 2025-12-10 cs.CV 50%

Scale-invariant and View-relational Representation Learning for Full Surround Monocular Depth

尺度不变与视图关系的表示学习用于全环绕单目深度估计

Kyumin Hwang, Wonhyeok Choi, Kiljoon Han, Wonjoon Choi, Minwoo Choi, Yongcheon Na, Minwoo Park, Sunghoon Im

机构 * Department of Electrical Engineering & Computer Sciences, Daegu Gyeongbuk Institute of Science and Technology (DGIST)(电子工程与计算机科学系,大邱庆北科学技术大学) Department of Autonomous Driving Perception Technology Vanguard Team, Hyundai Motor Company(自动驾驶感知技术先锋团队,现代汽车公司)

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 本文提出了一种结合尺度不变与视图关系的知识蒸馏方法,用于提升全环绕单目深度估计的性能与效率。

Comments Accepted at IEEE Robotics and Automation Letters (RA-L) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06387 2025-12-09 cs.CR cs.RO 50%

Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence

超越模型劫持:具身智能中的“十大致命罪恶”系统性分析

Yuhang Huang, Junchao Li, Boyang Ma, Xuelong Dai, Minghui Xu, Kaidi Xu, Yue Zhang, Jianping Wang, Xiuzhen Cheng

机构 * Shandong University(山东大学) City University of Hong Kong(香港城市大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本研究首次系统分析了具身智能平台的十大安全漏洞,揭示了跨层安全弱点,提出了构建稳健具身系统的建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05677 2025-12-08 stat.ME math.PR stat.ML 50%

Empirical Decision Theory

经验决策理论

Christoph Jansen, Georg Schollmeyer, Thomas Augustin, Julian Rodemann

专题命中 知识编辑与模型理解 :prompting(abstract)

AI总结 本文提出了一种经验决策模型,通过协议中的观察行动-后果对来处理决策问题,无需显式指定世界状态,提供了三种推断保证方法。

Comments Christoph Jansen and Georg Schollmeyer contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05482 2025-12-08 cs.CV 50%

Concept-based Explainable Data Mining with VLM for 3D Detection

基于概念的可解释数据挖掘与VLM用于3D检测

Mai Tsujimoto

机构 * The University of Tokyo(东京大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出基于概念的可解释数据挖掘方法,利用VLMs识别稀有物体以提升3D检测性能,减少标注负担并提高模型效果。

Comments 28 pages including appendix. Code: https://github.com/mm1129/concept_based_rare_detector_2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05349 2025-12-08 cond-mat.mtrl-sci 50%

Platonic representation of foundation machine learning interatomic potentials

柏拉图表示法用于基础机器学习互原子势

Zhenzhu Li, Aron Walsh

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 本研究提出柏拉图表示法,通过统一不同MLIPs的潜在空间,实现跨模型最优传输和可解释的嵌入运算,揭示潜在空间中几何失真与物理预测失败的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22345 2025-12-05 cs.CV 50%

Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment

反向流动:通过反向表征对齐改进规范化流

Yang Chen, Xiaowei Xu, Shuai Wang, Chenhui Zhu, Ruxue Wen, Xubin Li, Tiezheng Ge, Limin Wang

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 通过反向表征对齐改进规范化流,提升生成质量和分类准确性,训练速度提升3.3倍,实现ImageNet新state-of-the-art结果。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04483 2025-12-05 cs.CV 50%

DeRA: Decoupled Representation Alignment for Video Tokenization

DeRA: 解耦表示对齐用于视频标记化

Pengbo Guo, Junke Wang, Zhen Xing, Chengxu Liu, Daoguo Dong, Xueming Qian, Zuxuan Wu

机构 * Xi’an Jiaotong University(西安交通大学) Shanghai Innovation Institute(上海创新研究院) Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院)

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 DeRA通过解耦时空表示学习,提升视频标记化效率和性能,并在视频生成任务中取得新突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03601 2025-12-04 cs.CV 50%

Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding

Motion4D: 学习3D一致的运动和语义以实现4D场景理解

Haoran Zhou, Gim Hee Lee

专题命中 知识编辑与模型理解 :foundation model(abstract)

AI总结 Motion4D通过整合2D先验和4D高斯点撒表示,提升3D一致性和语义一致性,实现更准确的4D场景理解。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03577 2025-12-04 cs.CV 50%

Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning

跨染色对比学习用于配对免疫组化和病理切片的表示学习

Yizhi Zhang, Lei Fan, Zhulin Tao, Donglin Di, Yang Song, Sidong Liu, Cong Cong

机构 * Communication University of China(通信大学) Tsinghua University(清华大学) Macquarie University(麦考瑞大学)

专题命中 知识编辑与模型理解 :pretraining(abstract)

AI总结 本文提出CSCL方法,通过跨染色对比学习提升H&E与IHC特征的兼容性,实现高质量的切片级表示学习。

Comments 6 pages, 2 figures. Camera-ready version accepted for IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03142 2025-12-04 cs.CV 50%

UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying

UniEdit-I: 基于迭代理解、编辑和验证的无训练图像编辑

Chengyu Bai, Jintao Chen, Xiang Bai, Yilong Chen, Qi She, Ming Lu, Shanghang Zhang

机构 * Peking University(北京大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 UniEdit-I通过在语义潜在空间中引入迭代理解、编辑和验证循环,实现了无训练的闭环图像编辑,无需微调或架构修改即可达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03079 2025-12-03 cs.CV 50%

Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA

偏见超越人口统计:通过反事实视觉问答探测黑盒大视觉-语言模型的决策边界

Zaiying Zhao, Toshihiko Yamasaki

机构 * The University of Tokyo(东京大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文通过反事实VQA基准探测黑盒LVLMs的决策边界,揭示非人口属性对决策的更大影响,并展示人类规范验证示例对提升模型响应一致性和公平性的作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00557 2025-12-02 cs.CV 50%

NeuroVolve: Evolving Visual Stimuli toward Programmable Neural Objectives

NeuroVolve:通过可编程神经目标演化视觉刺激

Haomiao Chen, Keith W Jamison, Mert R. Sabuncu, Amy Kuceyeski

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技) Weill Cornell Medicine(韦尔医学院)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 NeuroVolve通过可编程神经目标生成视觉刺激,揭示大脑区域间的协同与对抗性调节关系,实现脑引导的图像编辑与首选刺激生成的统一。

详情

展开后加载摘要…

URL PDF HTML 收藏