arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-08 至 2025-12-08 共收录 139 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 23 篇

2512.05686 2025-12-08 eess.SY cs.SY 75%

LA-RL: Language Action-guided Reinforcement Learning with Safety Guarantees for Autonomous Highway Driving

LA-RL: 基于语言动作引导的安全强化学习用于自动驾驶高速公路驾驶

Yiming Shu, Jiahui Xu, Jiwei Tang, Ruiyang Gao, Chen Sun

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 LA-RL通过整合大语言模型的语义推理和改进的安全层,提升自动驾驶高速公路驾驶的效率与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23471 2025-12-08 cs.SE 75%

Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency

为评估和提升代码效率合成性能约束

Jun Yang, Cheng-Chi Wang, Bogdan Alexandru Stoica, Kexin Pei

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 WEDGE通过生成性能压力输入,提升代码优化效果,释放PERFFORGE测试以评估未来高效代码生成方法。

Comments Accepted by Neurips 2025 (main poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14779 2025-12-08 cs.CL cs.AI cs.LG 75%

Towards Data-efficient Customer Intent Recognition with Prompt-based Learning Paradigm

基于提示学习范式的高效客户意图识别

Hengyu Luo, Peng Liu, Stefan Esping

机构 * University of Helsinki(赫尔辛基大学) Ingka Group, IKEA(英格卡集团)

专题命中 效率与部署 :language model(abstract);small language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于提示学习的高效客户意图识别方法,通过减少数据依赖提升小型语言模型的识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10809 2025-12-08 cs.LG cs.AI 73%

Rethinking Sparse Autoencoders: Select-and-Project for Fairness and Control from Encoder Features Alone

重新思考稀疏自编码器:仅从编码器特征中进行选择和投影以实现公平性与控制

Antonio Bărbălau, Cristian Daniel Păduraru, Teodor Poncu, Alexandru Tifrea, Elena Burceanu

机构 * Bitdefender University Politehnica of Bucharest(巴特亚大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于编码器特征的选择和投影框架,通过改进公平性和可控性,提升模型在视觉语言和大语言模型中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18973 2025-12-08 cs.CL cs.LG 73%

Hierarchical Mamba Meets Hyperbolic Geometry: A New Paradigm for Structured Language Embeddings

层次Mamba与双曲几何:一种新的结构语言嵌入范式

Sarang Patil, Ashish Parmanand Pandey, Ioannis Koutis, Mengjia Xu

机构 * New Jersey Institute of Technology(新泽西理工学院)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出层次Mamba结合双曲几何,以学习具有层次意识的语言嵌入,提升复杂层次推理能力。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05916 2025-12-08 cs.LG 70%

KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity

KQ-SVD:通过可证明的注意力保真度保障压缩KV缓存

Damien Lesens, Beheshteh T. Rakhshan, Guillaume Rabusseau

机构 * ENS de Lyon(里昂高等师范学院) DIRO, Université de Montréal(蒙特利尔大学DIRO中心) Mila(Mila人工智能研究所) CIFAR AI Chair(CIFAR人工智能主席)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 KQ-SVD通过最优低秩分解提升KV缓存压缩的注意力保真度,实验证明其在LLaMA和Mistral模型中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05428 2025-12-08 cs.SE 67%

Bita: A Conversational Assistant for Fairness Testing

Bita:一个用于公平性测试的对话助手

Keeryn Johnson, Cleyton Magalhaes, Ronnie de Souza Santos

专题命中 效率与部署 :large language model(abstract);language model(abstract)

AI总结 Bita是一个基于大型语言模型的对话助手,旨在通过公平性视角帮助软件测试人员检测偏见、评估测试计划并生成公平性导向的测试章程,提供可重复的公平性测试解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04555 2025-12-08 cs.RO cs.CV 67%

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

Evo-1:轻量级视觉-语言-动作模型,保持语义对齐

Tao Lin, Yilei Zhong, Yuxin Du, Jingjing Zhang, Jiting Liu, Yinxinyu Chen, Encheng Gu, Ziyan Liu, Hongyi Cai, Yanwen Zou, Lixing Zou, Zhaoye Zhou, Gen Li, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) EvoMind Tech(EvoMind科技) IAAR-Shanghai(IAAR-上海) SII Carnegie Mellon University(卡内基梅隆大学) University of Cambridge(剑桥大学) Nanyang Technological University(南洋理工大学)

专题命中 效率与部署 :language model(abstract);pretraining(abstract)

AI总结 Evo-1是一种轻量级的视觉-语言-动作模型,通过减少计算并保持语义对齐,实现了高效的部署和强大的性能。

Comments Github: https://github.com/MINT-SJTU/Evo-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15436 2025-12-08 cs.CV 67%

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

基于动态视觉搜索和缩放的自适应聚焦推理方法用于高效VLMs

Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li

机构 * organization= School of Computer Science \& Technology, Beijing Institute of Technology , city= Beijing , country= China organization= State Key Laboratory of General Artificial Intelligence, BIGAI , city= Beijing , country= China organization= School of Intelligence Science Technology, Peking University , city= Beijing , country= China organization= Guangdong Laboratory of Machine Perception Intelligent Computing, Shenzhen MSU--BIT University , city= Shenzhen , country= China organization= Department of Automation, Tsinghua University , city= Beijing , country= China

专题命中 效率与部署 :language model(abstract);SFT(abstract)

AI总结 本文提出基于动态视觉搜索和缩放的自适应聚焦推理方法,提升VLMs的多模态推理效率和实际应用效果。

Comments https://github.com/xtong-zhang/Chain-of-Focus

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05511 2025-12-08 cs.CV 50%

Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm

重新思考红外小目标检测:一种基础驱动的高效范式

Chuang Yu, Jinmiao Zhao, Yunpeng Liu, Yaokun Li, Xiujun Shu, Yuanhao Feng, Bo Wang, Yimian Dai, Xiangyu Yue

机构 * Key Laboratory of Opto-Electronic Information Processing, Chinese Academy of Sciences(光电信息处理重点实验室,中国科学院) Shenyang Institute of Automation, Chinese Academy of Sciences(沈阳自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Sun Yat-sen University(中山大学) Tencent(腾讯) Nankai University(南开大学) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 效率与部署 :foundation model(abstract)

AI总结 本文提出FDEP范式,通过基础模型冻结表示与任务特征融合,提升红外小目标检测精度,构建综合评估指标,实现多数据集SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07223 2025-12-08 physics.ins-det astro-ph.IM nucl-ex 50%

Readout noise of digital frequency multiplexed TES detectors for CUPID

CUPID数字频率复用TES探测器的读出噪声

Michel Adamič, Joseph Camilleri, Chiara Capelli, Matt Dobbs, Tucker Elleflot, Yury G. Kolomensky, Daniel Mayer, Joshua Montgomery, Valentine Novosad, Vivek Singh, Graeme Smecher, Aritoki Suzuki, Bradford Welliver

专题命中 效率与部署 :prompting(abstract)

AI总结 CUPID实验采用新型高频复用TES读出系统,通过优化SQUID参数和减少导线电容以降低读出噪声,满足高能分辨要求。

Comments 8 pages, 8 figures. Presented at the IEEE SORMA West 2025, accepted for publication in IEEE Transactions on Nuclear Science

Journal ref IEEE Transactions on Nuclear Science

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 领域大模型 11 篇

2512.05863 2025-12-08 cs.CL cs.AI 90%

Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework

优化医疗问答系统:基于RAG框架的微调与零样本大语言模型比较研究

Tasnimul Hassan, Md Faisal Karim, Haziq Jeelani, Elham Behnam, Robert Green, Fayeq Jeelani Syed

机构 * Department of Electrical Engineering Computer Science University of Toledo Toledo, USA Institute of Mathematical Sciences Claremont Graduate University Claremont, USA Department of Bioengineering University of Toledo Toledo, USA Department of Computer Science Bowling Green State University Bowling Green, USA

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文通过RAG框架结合微调与零样本大语言模型,提升医疗问答系统的准确性与可靠性,实验证明检索增强显著提高回答质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05167 2025-12-08 cs.AI 89%

Bridging Traditional Machine Learning and Large Language Models: A Two-Part Course Design for Modern AI Education

连接传统机器学习与大语言模型:一种面向现代人工智能教育的双阶段课程设计

Fang Li

机构 * Computer Science Department(计算机科学系) Oklahoma Christian University(俄克拉荷马基督教大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出一种双阶段课程设计,旨在连接传统机器学习与大语言模型,帮助学生全面理解人工智能发展并提升实践能力。

Comments Accepted by the 39th annual Consortium for Computing Sciences in Colleges (CCSC:SE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05311 2025-12-08 cs.LG cs.AI 81%

The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?

LLM签名的消解:在迭代改写后能否仍区分人类和LLM生成的科学想法?

Sadat Shahriar, Navid Ayoobi, Arjun Mukherjee

机构 * University of Houston(德克萨斯大学)

专题命中 领域大模型 :LLM(title,abstract);分类 cs.AI、cs.LG

AI总结 研究探讨了在多次迭代改写后,现有模型区分人类与LLM生成科学想法的能力,发现改写显著削弱了LLM签名的可区分性。

Comments Published in RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15432 2025-12-08 cs.CV 78%

MedDiff-FM: A Diffusion-based Foundation Model for Versatile Medical Image Applications

MedDiff-FM: 一种基于扩散的多功能医学图像应用基础模型

Yongrui Yu, Yannian Gu, Shaoting Zhang, Xiaofan Zhang

机构 * Qing Yuan Research Institute, Shanghai Jiao Tong University(上海交通大学清元研究 institute)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 MedDiff-FM是一种基于扩散的医学图像基础模型,通过预训练和微调实现多种医学图像处理任务,如去噪、异常检测和超分辨率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04535 2025-12-08 cs.AI 77%

GTM: Simulating the World of Tools for AI Agents

GTM: 为AI代理模拟工具世界

Zhenzhen Ren, Xinpeng Zhang, Zhenxing Qian, Yan Gao, Yu Shi, Shuxin Zheng, Jiyan He

机构 * Zhongguancun Academy, Beijing, China(中关村学院,北京,中国) School of Computer Science, Fudan University, Shanghai, China(计算机学院,复旦大学,上海,中国)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 GTM通过模拟工具提升AI代理训练效率,实现快速且低成本的工具交互模拟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01788 2025-12-08 cs.HC cs.CY 75%

Exploring ChatGPT's Capabilities, Stability, Potential and Risks in Conducting Psychological Counseling through Simulations in School Counseling

探索ChatGPT在模拟学校咨询中的能力、稳定性、潜力与风险

Yang Ni, Yanzhuo Cao

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本研究评估ChatGPT在模拟学校咨询中的应用,探讨其作为培训工具的潜力及伦理风险,提出使用指南。

Journal ref Mental Health and Digital Technologies, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05858 2025-12-08 cs.CL 74%

Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy

提示科学报告4:扮演角色:专家角色不提高事实准确性

Savir Basil, Ina Shapiro, Dan Shapiro, Ethan Mollick, Lilach Mollick, Lennart Meincke

专题命中 领域大模型 :prompting(title);分类 cs.CL

AI总结 该研究探讨了专家角色和低知识角色对AI模型在客观多项选择题上的准确性影响,发现角色提示通常未提升表现,专家角色在多数情况下无益,低知识角色反而降低准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17948 2025-12-08 cs.IR cs.AI 70%

VERIRAG: A Post-Retrieval Auditing of Scientific Study Summaries

VERIRAG:一种科学研究摘要的后检索审计

Shubham Mohole, Hongjun Choi, Shusen Liu, Christine Klymko, Shashank Kushwaha, Derek Shi, Wesam Sakla, Sainyam Galhotra, Ruben Glatt

机构 * Cornell University(康奈尔大学) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) University of Illinois Urbana‑Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 领域大模型 :language model(abstract);small language model(abstract);分类 cs.AI

AI总结 VERIRAG通过后检索审计框架检测科学摘要的方法学漏洞,提升宏F1值19个百分点,提供结构化审计轨迹支持负责任的科学实践。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15721 2025-12-08 cs.LG stat.ML 70%

Privacy-Preserving Conformal Prediction Under Local Differential Privacy

在局部差分隐私下实现隐私保护的置信区间

Coby Penso, Bar Mahpud, Jacob Goldberger, Or Sheffet

机构 * Bar-Ilan University(巴伊兰大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出在局部差分隐私下实现隐私保护的置信区间方法,通过两种互补的隐私保护技术,确保在不可信聚合者环境下仍能保持数据和标签隐私,并提供稳健的覆盖保证。

Comments Accepted by COPA 2025. Proceedings of Machine Learning Research 26, 2025 Conformal and Probabilistic Prediction with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01309 2025-12-08 cs.DB cs.AI 70%

Enhancing SPARQL Query Rewriting for Complex Ontology Alignments

增强SPARQL查询重写以应对复杂本体对齐

Anicet Lepetit Ondo, Laurence Capus, Mamadou Bousso

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出基于自然语言需求的SPARQL查询重写方法,利用等价传递性与大语言模型处理复杂本体对齐,提升查询效率与用户可访问性。

Comments This update corrects a minor error in Table 1 of the originally submitted version, where a formula was inadvertently included. This does not affect the methodology or results. We also improved the formatting of existing formulas and the visual presentation of algorithm outputs. No changes were made to the scientific content

Journal ref International Journal of Web & Semantic Technology (IJWesT) Vol.16, No.2, April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05179 2025-12-08 cs.CL cs.AI 62%

Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale

针对特定领域的问题回答微调BERT:迈向大学规模的自然语言处理资源

Aurélie Montfrond

机构 * Dept. of Electronic & Computer Engineering, University of Limerick(电子与计算机工程系,利默里克大学)

专题命中 领域大模型 :foundation model(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过微调BERT模型,针对大学课程材料开发了首个领域特定问答系统,展示了基础模型在教育领域的应用潜力。

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 知识编辑与模型理解 12 篇

2512.05461 2025-12-08 cs.CY cs.AI cs.HC 85%

Knowing Your Uncertainty -- On the application of LLM in social sciences

了解你的不确定性——关于LLM在社会科学中的应用

Bolun Zhang, Linzhuo Li, Yunqi Chen, Qinlin Zhao, Zihan Zhu, Xiaoyuan Yi, Xing Xie

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一个统一框架,用于评估LLM在社会科学任务中的不确定性,通过任务类型和验证类型两个维度,为研究者提供方法学保障和实践指导。

Comments 49 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16188 2025-12-08 cs.CL 83%

SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models

SAE-SSV: 在稀疏表示空间中进行监督引导以可靠控制语言模型

Zirui He, Mingyu Jin, Bo Shen, Ali Payani, Yongfeng Zhang, Mengnan Du

机构 * NJIT(新 jersey 理工学院) Rutgers University(罗格斯大学) Cisco(思科公司)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 SAE-SSV通过在稀疏表示空间中训练监督引导向量,实现对语言模型行为的可靠控制,提高了生成任务的成功率和可解释性。

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05546 2025-12-08 cs.CV cs.AI 79%

Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models

有意识的注视:用于视觉-语言模型中幻觉抑制的自适应注意力机制

Weijue Bu, Guan Yuan, Guixian Zhang

机构 * School of Computer Science and Technology/School of Artificial Intelligence(计算机科学与技术学院/人工智能学院) China University of Mining and Technology(中国矿业大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 CG-VLM通过认知需求传感器和聚焦共识诱导模块,在推理时精准干预视觉-语言模型的注意力,有效抑制幻觉并提升性能。

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00713 2025-12-08 cs.CR cs.AI 79%

Concept-Guided Backdoor Attack on Vision Language Models

基于概念的视觉语言模型后门攻击

Haoyu Shen, Weimin Lyu, Haotian Xu, Tengfei Ma

机构 * Stony Brook University(石溪大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本研究提出基于概念的视觉语言模型后门攻击方法,通过概念阈值污染和概念瓶颈模型引导未见后门两种技术,实现对模型生成文本的恶意替换。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19532 2025-12-08 q-bio.BM 78%

Toward the Explainability of Protein Language Models

迈向蛋白质语言模型的可解释性

Andrea Hunklinger, Noelia Ferruz

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文探讨了XAI在蛋白质语言模型中的应用,提出了XAI在蛋白质研究中的五个潜在角色,并呼吁推动可解释性的发展。

Comments 15 pages, 6 figures; version 4: Additional revision of the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17162 2025-12-08 cs.CR 78%

Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model

像二进制一样分析PDF:通过中间表示和语言模型实现对抗鲁棒的PDF恶意软件分析

Side Liu, Jiang Ming, Guodong Zhou, Xinyi Liu, Jianming Fu, Guojun Peng

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 通过中间表示和语言模型实现对抗鲁棒的PDF恶意软件分析,利用语义和结构特征提取提升检测性能。

Comments Accepted by ACM CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02327 2025-12-08 stat.ML cs.LG 77%

Variational Uncertainty Decomposition for In-Context Learning

变分不确定性分解用于上下文学习

I. Shavindra Jayasekera, Jacob Si, Filippo Valdettaro, Wenlong Chen, A. Aldo Faisal, Yingzhen Li

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出变分不确定性分解框架,用于分解上下文学习中的先验和随机不确定性,通过优化辅助查询获得不确定性上界并诱导下界,实验验证其有效性。

Comments Neurips Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05745 2025-12-08 cs.CR cs.MM 67%

ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior

ARGUS: 通过引导指令遵循行为防御多模态间接提示注入攻击

Weikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li, Huiping Zhuang, Ruidong Wang, Cen Chen, Hao Peng

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 ARGUS通过引导指令遵循行为,在表示空间中寻找最优防御方向,实现对多模态间接提示注入攻击的有效防御,同时保持模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏