arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-03-03 至 2026-03-03 共收录 599 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 124 篇

2510.22975 2026-03-03 cs.CV cs.GR cs.LG 57%

VoMP: Predicting Volumetric Mechanical Property Fields

VoMP:预测体积机械属性场

Rishit Dagli, Donglai Xiang, Vismay Modi, Charles Loop, Clement Fuji Tsang, Anka He Chen, Anita Hu, Gavriel State, David I. W. Levin, Maria Shugrina

机构 * NVIDIA University of Toronto(多伦多大学)

专题命中 评测与基准 :language model(abstract);分类 cs.LG

AI总结 VoMP通过训练几何变换器预测3D物体的体积机械属性场,结合多视角特征和现实世界数据集,实现高精度和高速度的属性估计。

Comments Project Page and hi-res paper: https://research.nvidia.com/labs/sil/projects/vomp

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23415 2026-03-03 cs.AI 57%

From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents

从对话到查询执行:为电子健康记录数据库代理构建用户和工具交互基准

Gyubok Lee, Woosog Chay, Heeyoung Kwak, Yeong Hwa Kim, Haanju Yoo, Oksoon Jeong, Meong Hi Son, Edward Choi

机构 * KAIST(韩国科学技术院) NAVER Cloud(NAVER云) Samsung Medical Center(三星医疗中心)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 EHR-ChatQA通过模拟环境评估数据库代理在处理用户查询模糊性和价值不匹配时的性能,揭示了构建安全关键EHR领域中鲁棒代理的重要性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14205 2026-03-03 cs.CL 57%

AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents

AgentSynth: 通用计算机使用代理的可扩展任务生成

Jingxu Xie, Dylan Xu, Xuandong Zhao, Dawn Song

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 AgentSynth通过生成多样化任务和轨迹数据集,提升通用计算机使用代理的训练效率和性能评估的准确性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12880 2026-03-03 cs.AI cs.MM 57%

Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey

多模态学习是否在医疗领域实现了通用智能?一项全面的综述

Qika Lin, Yifan Zhu, Xin Mei, Ling Huang, Jingying Ma, Kai He, Zhen Peng, Erik Cambria, Mengling Feng

机构 * Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学公共健康学院) School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院) School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

AI总结 本文通过全面调查,指出当前多模态学习在医疗领域尚未实现通用智能,并提出十个潜在研究方向。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00840 2026-03-03 cs.CL 57%

Learning Nested Named Entity Recognition from Flat Annotations

从平面标注中学习嵌套命名实体识别

Igor Rozhkov, Natalia Loukachevitch

机构 * Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 本文提出从平面标注中学习嵌套命名实体识别的方法,通过四种技术组合在NEREL基准上达到26.37%的内部F1,缩小了与完全嵌套监督的差距。

Comments Accepted at EACL 2026, 15 pages, 2 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00634 2026-03-03 cs.CL 57%

BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

BLUFF:跨58种低资源语言的虚假和合成内容检测基准测试

Jason Lucas, Matt Murtagh-White, Adaku Uchendu, Ali Al-Lawati, Michiharu Yamashita, Dominik Macko, Ivan Srba, Robert Moro, Dongwon Lee

机构 * Penn State University(宾夕法尼亚州立大学) Trinity College Dublin(都柏林圣三一学院) MIT Lincoln Lab(麻省理工学院林肯实验室) Visa Research(Visa研究) Kempelen Institute of Intelligent Technologies(凯普勒智能技术研究所)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 BLUFF通过覆盖58种低资源语言的多语言基准测试,提供检测虚假和合成内容的全面框架,结合人工与AI生成内容,推动公平的虚假信息检测研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00462 2026-03-03 cs.CV cs.AI 57%

OPGAgent: An Agent for Auditable Dental Panoramic X-ray Interpretation

OPGAgent: 一种用于可审计牙科全景X光解读的智能体

Zhaolin Yu, Litao Yang, Ben Babicka, Ming Hu, Jing Hao, Anthony Huang, James Huang, Yueming Jin, Jiasong Wu, Zongyuan Ge

机构 * AIM for Health Lab Faculty of Information Technology, Monash University(信息科技学院,莫纳什大学) Monash University(莫纳什大学) Curae Health Faculty of Dentistry, The University of Hong Kong(牙科学院,香港大学) National University of Singapore(新加坡国立大学) Southeast University(东南大学)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 OPGAgent是一种用于可审计牙科全景X光解读的智能体系统,通过多工具协调和共识机制,在结构化报告和VQA评估中优于现有牙科VLMs和医疗智能体框架。

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00434 2026-03-03 cs.ET cs.CL cs.IR 57%

RTLocating: Intent-aware RTL Localization for Hardware Design Iteration

RTLocating:面向硬件设计迭代的意图感知RTL定位

Changwen Xing, Yanfeng Lu, Lei Qi, Chenxu Niu, Jie Li, Xi Wang, Yong Chen, Jun Yang

机构 * School of Integrated Circuits, Southeast University, Nanjing, China(东南大学集成电路学院) National Center of Technology Innovation for EDA, Nanjing, China(EDA技术创新国家中心) School of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程学院) Department of Computer Science, Texas Tech University, Lubbock, USA(塔拉斯大学计算机科学系)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 RTLocating通过意图感知的RTL定位框架,实现了工业级硬件设计迭代中意图驱动的高效定位,显著提升了变更请求与RTL代码的匹配精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00432 2026-03-03 cs.CL 57%

A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs

基于语言类型学的多语言掩码语言模型中词序和形态敏感性评估框架

Anna Feldman, Libby Barak, Jing Peng

专题命中 评测与基准 :language model(abstract);分类 cs.CL

AI总结 本文提出了一种基于语言类型学的评估框架,用于评估多语言掩码语言模型对词序和形态敏感性的依赖性,通过多种扰动方法测试模型性能。

Journal ref LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00369 2026-03-03 cs.CL 57%

Policy Compliance of User Requests in Natural Language for AI Systems

人工智能系统中用户自然语言请求的政策合规性

Pedro Cisneros-Velarde

机构 * VMware Research(VMware研究院)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 本文提出首个评估AI系统用户请求政策合规性的基准,通过不同模型和方法评估合规性,揭示了该问题的挑战性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00285 2026-03-03 cs.AI 57%

TraderBench: How Robust Are AI Agents in Adversarial Capital Markets?

TraderBench: AI代理在对抗性资本市场中的鲁棒性如何?

Xiaochuang Yuan, Hui Xu, Silvia Xu, Cui Zou, Jing Xiong

机构 * Amazon.com Inc.(亚马逊公司) Stony Brook University(石溪大学) Stanford University(斯坦福大学) University of Oklahoma(俄克拉荷马大学) UC Santa Cruz(加州大学圣克鲁兹分校)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 TraderBench通过结合静态任务与对抗性交易模拟,评估AI代理在动态市场中的鲁棒性,发现现有模型在加密交易中表现稳定但缺乏真实适应性。

Comments Equal Contribution: Xiaochuang Yuan and Hui Xu contributed equally to this work. All correspondence should be directed to yxc20098@gmail.com. Submitted to Agents in the Wild Workshop, ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00206 2026-03-03 cs.CV cs.AI 57%

TACIT Benchmark: A Programmatic Visual Reasoning Benchmark for Generative and Discriminative Models

TACIT基准:一个生成和判别模型的程序化视觉推理基准

Daniel Nobrega Medeiros

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 TACIT基准提出一个程序化视觉推理测试,包含10个任务和6个领域,通过生成和判别双轨评估模型在空间导航、抽象模式完成等任务中的推理能力。

Comments 10 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00122 2026-03-03 cs.CV cs.AI cs.IR 57%

NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence

NovaLAD:一种快速、优化CPU的文档提取流水线用于生成式AI和数据智能

Aman Ulla

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 NovaLAD是一种基于CPU优化的快速文档提取系统,通过结合YOLO模型和规则分组技术,实现高效准确的文档解析,适用于生成式AI和数据智能领域。

Comments 17 pages, 10 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02076 2026-03-03 econ.GN cs.HC q-fin.EC 50%

When an AI Judges Your Work: The Hidden Costs of Algorithmic Assessment

当AI评判你的作品:算法评估的隐藏成本

David Almog, Lucas Lippman, Daniel Martin

专题命中 评测与基准 :LLM(abstract)

AI总结 本文研究了AI评估对工人行为的影响,发现工人在AI评估下产出更多但质量更低,并更倾向于使用外部工具,但工具使用并未解释产出差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02047 2026-03-03 cs.CV 50%

NICO-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Understanding the Nicotine Public Health Crisis

NICO-RAG:多模态超图检索增强生成用于理解尼古丁公共卫生危机

Manuel Serna-Aguilera, Raegan Anderes, Page Dobbs, Khoa Luu

机构 * University of Arkansas(亚拉巴马大学) University of Arkansas for Medical Sciences(亚拉巴马大学医学科学分校)

专题命中 评测与基准 :language model(abstract)

AI总结 NICO-RAG通过多模态超图检索增强生成,解决尼古丁公共卫生危机中的信息检索与生成问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11772 2026-03-03 cs.CV 50%

Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation

Seg2Track-SAM2:基于SAM2的多目标跟踪与分割

Diogo Mendonça, Tiago Barros, Cristiano Premebida, Urbano J. Nunes

机构 * University of Coimbra(科英布拉大学) Institute of Systems and Robotics(系统与机器人研究所) Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 评测与基准 :foundation model(abstract)

AI总结 Seg2Track-SAM2基于SAM2提出多目标跟踪与分割框架,通过整合预训练检测器与专用模块,实现跟踪初始化、数据关联和细化,提升身份一致性和内存效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01547 2026-03-03 cs.CV 50%

PathMoE: Interpretable Multimodal Interaction Experts for Pediatric Brain Tumor Classification

PathMoE:用于儿童脑肿瘤分类的可解释多模态交互专家

Jian Yu, Joakim Nguyen, Jinrui Fang, Awais Naeem, Zeyuan Cao, Sanjay Krishnan, Nicholas Konz, Tianlong Chen, Chandra Krishnan, Hairong Wang, Edward Castillo, Ying Ding, Ankita Shukla

机构 * University of Texas(德克萨斯大学) Dell Children’s Medical Center(德尔儿童医学中心) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Nevada, Reno(内华达大学里诺分校)

专题命中 评测与基准 :foundation model(abstract)

AI总结 PathMoE通过整合多模态信息提升儿童脑肿瘤分类性能,揭示了不同模态间的交互作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01324 2026-03-03 cs.CV 50%

Open-Vocabulary vs Supervised Learning Methods for Post-Disaster Visual Scene Understanding

开放词汇与监督学习方法在灾后视觉场景理解中的应用

Anna Michailidou, Georgios Angelidis, Vasileios Argyriou, Panagiotis Sarigiannidis, Georgios Th. Papadopoulos

机构 * Department of Informatics and Telematics, Harokopio University of Athens(信息与电信学系,雅典哈罗科波斯大学) Department of Networks and Digital Media, Kingston University(网络与数字媒体系,金士顿大学) Department of Electrical and Computer Engineering, University of Western Macedonia(电子与计算机工程系,西马其顿大学) Archimedes, Athena Research Center(阿基米德,雅典研究中心)

专题命中 评测与基准 :pretraining(abstract)

AI总结 本文比较了监督学习与开放词汇视觉模型在灾后场景理解中的表现,发现监督学习在特定条件下仍更可靠,尤其在处理小物体和精细边界时。

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01108 2026-03-03 cs.CV 50%

GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation

GroundedSurg: 一种多手术流程的语言条件手术工具分割基准

Tajamul Ashraf, Abrar Ul Riyaz, Wasif Tak, Tavaheed Tariq, Sonia Yadav, Moloud Abdar, Janibul Bashir

机构 * King Abdullah University of Science and Technology (KAUST)(卡奥尔大学科学与技术学院) Thapar Institute of Engineering and Technology(塔帕尔工程与技术学院) The University of Queensland(昆士兰大学) Gaash Research Lab, National Institute of Technology Srinagar(加什研究实验室,锡纳加尔国家理工学院)

专题命中 评测与基准 :language model(abstract)

AI总结 GroundedSurg提出了一种多手术流程的语言条件手术工具分割基准,通过实例级接地和空间注释评估视觉-语言模型在临床真实场景中的性能。

Comments https://github.com/gaash-lab/GroundedSurg

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07367 2026-03-03 cs.SD 50%

FOCAL: A Novel Benchmarking Technique for Multi-modal Agents

FOCAL:多模态代理的新型基准测试技术

Anupam Purwar, Aditya Choudhary

专题命中 评测与基准 :language model(abstract)

AI总结 FOCAL提出了一种新型基准测试技术,用于评估多模态代理的端到端推理能力、误差传播及对话效果。

Comments We present a framework for evaluation of Multi-modal Agents consisting of Voice-to-voice model components viz. Text to Speech (TTS), Retrieval Augmented Generation (RAG) and Speech-to-text (STT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01029 2026-03-03 cs.CV 50%

Vision-Language Feature Alignment for Road Anomaly Segmentation

视觉-语言特征对齐用于道路异常分割

Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue

机构 * School of Computer Science, Fudan University(复旦大学计算机科学学院) Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑启发智能科学与技术研究院)

专题命中 评测与基准 :language model(abstract)

AI总结 VL-Anomaly通过结合视觉-语言模型的语义先验,提出了一种新的道路异常分割框架,有效提升异常检测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00990 2026-03-03 cs.CV 50%

MLRecon: Robust Markerless Freehand 3D Ultrasound Reconstruction via Coarse-to-Fine Pose Estimation

MLRecon: 通过粗到细的姿态估计实现鲁棒的无标记自由手3D超声重建

Yi Zhang, Puxun Tu, Kun Wang, Yulin Yan, Tao Ying, Xiaojun Chen

机构 * Institute of Biomedical Manufacturing and Life Quality Engineering, School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China(生物医学制造与生命质量工程学院,机械工程学院,上海交通大学,上海,中国) Department of Ultrasound in Medicine, Shanghai Sixth People's Hospital Affiliated to Shanghai Jiao Tong University School of Medicine, Shanghai, China(医学超声科,上海第六人民医院(隶属于上海交通大学医学院)) Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China(医学机器人研究院,上海交通大学,上海,中国)

专题命中 评测与基准 :foundation model(abstract)

AI总结 MLRecon通过粗到细姿态估计实现无标记自由手3D超声重建,有效解决漂移问题,实现高精度的6D探头跟踪与高质量3D成像。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00805 2026-03-03 cs.CV cs.MA 50%

NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code

NERFIFY:一个将NeRF论文转化为代码的多智能体框架

Seemandhar Jain, Keshav Gupta, Kunal Gupta, Manmohan Chandraker

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 评测与基准 :LLM(abstract)

AI总结 NERFIFY通过多智能体框架将NeRF论文转化为可训练的代码,提升复现效率和质量。

Comments Accepted to CVPR 2026. Project page: https://seemandhar.github.io/NERFIFY/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00080 2026-03-03 cs.DL 50%

From Static Repositories to Agentic Knowledge Webs: ResearchTwin and the S-Index for Federated Human-AI Research Discovery

从静态仓库到代理知识网络:ResearchTwin和S-指数用于联邦人机研究发现

Martin G. Frasch

专题命中 评测与基准 :LLM(abstract)

AI总结 ResearchTwin通过S-指数量化多模态研究影响,利用联邦架构实现人机协同的知识发现。

Comments 15 pages, 1 figure, https://github.com/martinfrasch/ResearchTwin

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 效率与部署 107 篇

2603.00846 2026-03-03 cs.IR cs.LG 90%

Tiny-Critic RAG: Empowering Agentic Fallback with Parameter-Efficient Small Language Models

Tiny-Critic RAG:赋能代理回退的参数高效小型语言模型

Yichao Wu, Penghao Liang, Yafei Xiang, Mengwei Yuan, Jianan Liu, Jing Yang, Xianyou Li, Weiran Yan

机构 * Northeastern University(东北大学) Washington University in St. Louis(华盛顿大学圣路易斯分校) New York University(纽约大学)

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);large language model(abstract);SLM(abstract)

AI总结 Tiny-Critic RAG通过参数高效的Small Language Model实现低延迟的二进制路由,有效降低代理部署成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00196 2026-03-03 cs.CR cs.AI cs.CL 90%

Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models

您的推理请求将变成一个黑箱:面向云上大型语言模型的保密推理

Chung-ju Huang, Huiqiang Zhao, Yuanpeng He, Lijian Li, Wenpin Jiao, Zhi Jin, Peixuan Chen, Leye Wang

机构 * Key Lab of High Confidence Software Technologies (Peking University), Ministry of Education, China(高可信软件技术重点实验室(北京大学)) School of Computer Science, Peking University, Beijing, China(北京大学计算机学院) Tencent, Shenzhen, China(腾讯(深圳)) Macau university, China(澳门大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 Talaria通过在客户端控制的CVM中执行敏感操作并使用ReMO协议保护隐私,实现云上LLM的保密推理,同时保持性能和效率。

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01788 2026-03-03 cs.CL 89%

nchellwig at SemEval-2026 Task 3: Self-Consistent Structured Generation (SCSG) for Dimensional Aspect-Based Sentiment Analysis using Large Language Models

nchellwig 在 SemEval-2026 任务 3: 使用大语言模型进行维度方面基于情感分析的自一致结构生成(SCSG)

Nils Constantin Hellwig, Jakob Fehle, Udo Kruschwitz, Christian Wolff

机构 * Media Informatics Group, University of Regensburg(雷姆施泰因大学媒体信息学小组) Information Science Group, University of Regensburg(雷姆施泰因大学信息科学小组)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 nchellwig 在 SemEval-2026 任务 3 中提出 SCSG 方法,利用大语言模型进行维度方面基于情感分析,通过自一致性生成提升预测可靠性,实验显示在多个语言和领域组合中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00694 2026-03-03 cs.RO cs.AI 89%

Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model

Wild-Drive: 通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划

Zihang Wang, Xu Li, Benwu Wang, Wenkai Zhu, Xieyuanli Chen, Dong Kong, Kailin Lyu, Yinan Du, Yiming Peng, Haoyang Che

机构 * School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院) Southeast University Nanjing Jiangbei New Area Innovation Research Institute(东南大学南京江滨新区创新研究院) National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology(国防科技大学装备状态感知与智能支撑国家重点实验室) School of Transportation, Shandong University of Science and Technology(山东科技大学交通学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 Wild-Drive通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划,提升复杂环境下的可解释性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00510 2026-03-03 cs.CV cs.AI 89%

What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models

视觉标记究竟编码了什么?揭示多模态大语言模型中的稀疏性与冗余性

Yingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本研究揭示多模态大语言模型中视觉标记的稀疏性与冗余性,通过EmbedLens工具发现活跃标记在进入模型前已编码细粒度信息,并提出中层注入方法提升效率。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03772 2026-03-03 cs.AI 89%

A Contemporary Overview: Trends and Applications of Large Language Models on Mobile Devices

大型语言模型在移动设备上的趋势与应用综述

Lianjun Liu, Hongli An, Pengxuan Chen, Longxiang Ye

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文综述了大型语言模型在移动设备上的发展趋势与应用,探讨了其在语音助手、实时翻译等领域的潜力及对智能设备和物联网的推动作用。

Comments The authors withdraw this manuscript. A substantially revised version will be submitted later

详情

展开后加载摘要…

URL PDF HTML 收藏