arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-02 至 2025-12-02 共收录 23 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 23 篇

2512.00651 2025-12-02 cs.SE cs.LG 89%

Large Language Models for Software Engineering: A Reproducibility Crisis

大型语言模型在软件工程中的应用:可重复性危机

Mohammed Latif Siddiq, Arvin Islam-Gomes, Natalie Sekerak, Joanna C. S. Santos

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG

AI总结 本文探讨了基于大型语言模型的软件工程研究中可重复性问题,揭示了制品可用性、环境规范和文档清晰度等方面的不足,并提出了可重复性成熟度模型以提升评估的全面性。

Comments Submitted to Empirical Software Engineering (EMSE) journal; 112 pages (81 pages of references)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01431 2025-12-02 cs.HC 89%

A Meta-Analysis of the Persuasive Power of Large Language Models

大语言模型说服力的元分析

Lukas Hölbling, Sebastian Maier, Stefan Feuerriegel

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本研究通过元分析发现,大语言模型在说服效果上与人类无显著差异,但情境因素如模型、对话设计和领域对说服表现有重要影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01330 2025-12-02 cs.DL 89%

Prompt perturbation and fraction facilitation sometimes strengthen Large Language Model scores

提示扰动与分数促进有时会增强大语言模型得分

Mike Thelwall

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 本文研究了如何通过提示扰动和分数促进策略提升LLM对期刊文章质量评分的准确性,发现平均语义等价提示和允许分数评分有助于提高模型表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01191 2025-12-02 cs.CL 88%

Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks

通用大语言模型在医疗基准测试中优于临床工具

Krithik Vishwanath, Mrigayu Ghosh, Anton Alyakin, Daniel Alexander Alber, Yindalon Aphinyanaphongs, Eric Karl Oermann

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 通用大语言模型在医疗基准测试中表现优于临床工具,揭示了临床AI系统在某些关键能力上的不足。

Comments 17 pages, 4 figures (2 regular, 2 supplemental)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16237 2025-12-02 cs.IR cs.LG 87%

LLM-Enhanced Reranking for Complementary Product Recommendation

基于大语言模型的互补产品推荐增强重排序

Zekun Xu, Yudi Zhang

机构 * North Carolina State University(北卡罗来纳州立大学) Iowa State University(爱荷华州立大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出利用大语言模型增强互补产品推荐的重排序,通过直接应用LLM提示策略提升推荐准确性和多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01893 2025-12-02 cs.CR 85%

Improving Phishing Resilience with AI-Generated Training: Evidence on Prompting, Personalization, and Duration

利用AI生成训练提升钓鱼攻击抵御能力:关于提示、个性化和持续时间的证据

Francesco Greco, Giuseppe Desolda, Cesare Tucci, Andrea Esposito, Antonio Curci, Antonio Piccinno

专题命中 领域大模型 :prompting(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文通过实验验证了利用大型语言模型生成钓鱼攻击抵御培训的有效性,发现简单提示策略即可产生显著学习提升,且无需复杂个性化。

Comments Data and code available at: https://doi.org/10.6084/m9.figshare.30664793

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00597 2025-12-02 cs.CV 85%

Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models

缩小规模以扩大规模:通过跨模态低秩适应实现操作高效且可部署的临床模型

Thuraya Alzubaidi, Farhad R. Nezami, Muzammil Behzad

机构 * King Fahd University of Petroleum(国王法赫德石油与矿物大学) Institute for Medical Engineering(医学工程研究所) Science, Massachusetts Institute of Technology, US(科学,麻省理工学院,美国) Harvard Medical School, Harvard University, US(哈佛医学院,哈佛大学,美国) SDAIA-KFUPM Joint Research Center for Artificial Intelligence, Saudi Arabia(SDAIA-KFUPM人工智能联合研究中心,沙特阿拉伯)

专题命中 领域大模型 :language model(title,abstract);foundation model(abstract);pretraining(abstract)

AI总结 通过跨模态低秩适应,MedCT-VLM在零样本分类中实现了对CT影像的高效适应,显著提升了病理分类的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10044 2025-12-02 cs.CR cs.AI 84%

Large Language Models for Power System Security: A Novel Multi-Modal Approach for Anomaly Detection in Energy Management Systems

用于电力系统安全的大型语言模型:一种用于能源管理系统异常检测的新型多模态方法

Aydin Zaboli, Junho Hong, Alexandru Stefanov, Chen-Ching Liu, Chul-Sang Hwang

机构 * Department of Electrical and Computer Engineering, University of Michigan -- Dearborn, MI, 48128 USA.(电气与计算机工程系,密歇根大学迪尔伯恩分校) Department of Electrical Sustainable Energy, Technische Universiteit Delft, 2628 CD Delft, Netherlands.(可持续能源系,代尔夫特理工大学) Bradley Department of Electrical and Computer Engineering, Virginia Polytechnic Institute and State University, Blacksburg, VA 24061, USA.(布雷德利电气与计算机工程系,弗吉尼亚理工学院和州立大学) Smart Grid Research Division System Reliability Research Team, Korea Electrotechnology Research Institute (KERI), Gwangju-si, 61751, South Korea.(智能电网研究分会系统可靠性研究团队,韩国电力技术研究所(KERI))

专题命中 领域大模型 :large language model(title);language model(title);分类 cs.AI

AI总结 本文提出了一种基于大型语言模型的多模态方法,用于电力系统中能源管理系统的安全防护和异常检测。

Comments 10 Figures; 6 Tables; Accepted, IEEE ACCESS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01453 2025-12-02 q-bio.OT 82%

Reinventing Clinical Dialogue: Agentic Paradigms for LLM Enabled Healthcare Communication

重新定义临床对话:面向LLM赋能医疗沟通的代理范式

Xiaoquan Zhi, Hongke Zhao, Likang Wu, Chuang Zhao, Hengshu Zhu

专题命中 领域大模型 :LLM(title);large language model(abstract);language model(abstract)

AI总结 本文提出一种新的分类法,通过知识源和代理目标的正交轴分析,揭示医疗AI从生成文本预测转向代理自主性的认知架构,并探讨四种范式在创造力与可靠性间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27680 2025-12-02 cs.CV cs.AI cs.LG 81%

PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting

PETAR:基于掩码感知的视觉-语言建模的局部发现生成用于PET自动报告

Danyal Maqbool, Changhee Lee, Zachary Huemann, Samuel D. Church, Matthew E. Larson, Scott B. Perlman, Tomas A. Romero, Joshua D. Warner, Meghan Lubner, Xin Tie, Jameson Merkow, Junjie Hu, Steve Y. Cho, Tyler J. Bradshaw

机构 * University of Wisconsin–Madison Department of Computer Sciences(威斯康星大学麦迪逊分校计算机科学系) University of Wisconsin–Madison Department Radiology(威斯康星大学麦迪逊分校放射学系) Microsoft(微软公司)

专题命中 领域大模型 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 PETAR通过引入PETARSeg-11K数据集和PETAR-4B模型,实现基于掩码感知的3D PET自动报告生成,提升医学影像分析的精度与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01434 2025-12-02 cs.AI 79%

A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building

一种灵活的多智能体LLM-人类框架,用于快速的人类验证工具构建

Daull Xavier, Patrice Bellot, Emmanuel Bruno, Vincent Martin, Elisabeth Murisasco

机构 * Toulon Univ(图卢兹大学) Aix Marseille Univ(阿维尼翁-马赛大学) CNRS(国家科学研究中心) LIS(信息系统实验室)

专题命中 领域大模型 :LLM(title,abstract);分类 cs.AI

AI总结 该研究提出了一种灵活的多智能体框架,通过人类反馈和强化学习,实现快速的人类验证工具构建,适用于复杂迭代任务。

Journal ref 2025 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), Nov 2025, Londres, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01010 2025-12-02 cs.MA cs.AI cs.LG cs.SE physics.comp-ph physics.flu-dyn 79%

Chain of Unit-Physics: A Primitive-Centric Approach to Scientific Code Synthesis

单元物理链:一种以基础原理为中心的科学代码合成方法

Vansh Sharma, Venkat Raman

机构 * University of Michigan(密歇根大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);RLHF(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出单元物理链框架,通过以基础原理为中心的多代理系统,有效解决科学代码生成中的可靠性问题,实现高精度和高效能的代码生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01922 2025-12-02 cs.CV 78%

Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding

Med-VCD: 通过视觉对比解码缓解医疗大视觉语言模型的幻觉问题

Zahra Mahdavi, Zahra Khodakaramimaghsoud, Hooman Khaloo, Sina Bakhshandeh Taleshani, Erfan Hashemi, Javad Mirzapour Kaleybar, Omid Nejati Manzari

机构 * Department of computer science, University of Central Florida, Orlando, USA(计算机科学系,中央佛罗里达大学) Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, USA(生物工程系,宾夕法尼亚大学) Department of electrical engineering, Columbia university, New York, NY, USA(电气工程系,哥伦比亚大学) Technical University of Applied Sciences Regensburg, Regensburg, Germany(应用科学技术大学(雷根斯堡)) Department of Surgery, University of Calgary, Calgary, Alberta, Canada(外科系,卡尔加里大学) University College of Nabi Akram, Tabriz, Iran(纳比阿克兰大学) School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran(电气工程学院,伊朗科学技术大学)

专题命中 领域大模型 :language model(title,abstract)

AI总结 Med-VCD通过视觉对比解码方法提升医疗大视觉语言模型的事实准确性与幻觉准确性,减少幻觉输出并提高推理效率。

Journal ref Computers in Biology and Medicine (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01214 2025-12-02 cs.CV cs.AI 77%

M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis

M4-BLIP:通过面部增强的局部分析推进多模态媒体篡改检测

Hang Wu, Ke Sun, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing(多媒体可信感知与高效计算重点实验室)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 M4-BLIP通过引入面部增强的局部分析,提升多模态媒体篡改检测的准确性和可解释性。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00596 2025-12-02 cs.IR cs.AI 77%

DLRREC: Denoising Latent Representations via Multi-Modal Knowledge Fusion in Deep Recommender Systems

DLRREC: 通过深度融合多模态知识在深度推荐系统中进行潜在表示去噪

Jiahao Tian, Zhenkai Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 DLRREC通过深度融合多模态和协同知识,提升深度推荐系统中潜在表示的去噪能力,从而实现更精确的推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11032 2025-12-02 cs.CL cs.CY 77%

DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4

DeID-GPT:通过GPT-4实现零样本医疗文本去标识化

Zhengliang Liu, Yue Huang, Xiaowei Yu, Lu Zhang, Zihao Wu, Chao Cao, Haixing Dai, Lin Zhao, Yiwei Li, Peng Shu, Fang Zeng, Lichao Sun, Wei Liu, Dinggang Shen, Quanzheng Li, Tianming Liu, Dajiang Zhu, Xiang Li

机构 * The University of Georgia(佐治亚大学) Lehigh University(莱恩大学) The University of Texas at Arlington(德克萨斯大学阿灵顿分校) Massachusetts General Hospital and Harvard Medical School(麻省总医院和哈佛医学院) Mayo Clinic(梅奥诊所) ShanghaiTech University(上海科技大学) Shanghai United Imaging Intelligence Co., Ltd.(上海联影智能科技有限公司) Shanghai Clinical Research and Trial Center(上海临床研究与试验中心)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 DeID-GPT利用GPT-4实现医疗文本的零样本去标识化,具有高准确性和可靠性,有效保护隐私信息。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01690 2025-12-02 cs.SE 75%

Generating REST API Tests With Descriptive Names

生成具有描述性名称的REST API测试

Philip Garrett, Juan P. Galeotti, Andrea Arcuri, Alexander Poth, Olsi Rrjolli

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出三种确定性技术生成REST API测试的描述性名称,并通过实验验证其有效性,证明轻量级方法可作为替代LLM方法的实用方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01224 2025-12-02 cs.LG 70%

CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions

CoSineVerifier:面向计算导向科学问题的工具增强答案验证工具

Ruixiang Feng, Zhenwei An, Yuntao Wen, Ran Le, Yiming Jia, Chen Yang, Zongchao Chen, Lisi Chen, Shen Gao, Shuo Shang, Yang Song, Tao Zhang

机构 * Nanbeige Lab, BOSS Zhipin(纳比实验室,BOSS智聘) University of Electronic Science and Technology of China(电子科技大学)

专题命中 领域大模型 :LLM(abstract);language model(abstract);分类 cs.LG

AI总结 CoSineVerifier通过工具增强的验证器提升科学问题的计算能力,实现超越语义匹配的验证效果,并在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00194 2025-12-02 cs.CV cs.LG eess.IV q-bio.QM 70%

AutocleanEEG ICVision: Automated ICA Artifact Classification Using Vision-Language AI

AutocleanEEG ICVision:利用视觉-语言AI实现自动化ICA噪声分类

Zag ElSayed, Grace Westerkamp, Gavin Gammoh, Yanchen Liu, Peyton Siekierski, Craig Erickson, Ernest Pedapati

机构 * School of Information Technology University of Cincinnati Ohio, USA(信息科技学院 奥地利辛辛那提大学 美国俄亥俄州)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 ICVision通过视觉-语言AI实现自动化EEG ICA噪声分类,达到专家级准确度并提供可解释结果。

Comments 6 pages, 8 figures

Journal ref Conference ICMI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00586 2025-12-02 cs.LG cs.CL q-bio.QM 62%

Statistical NLP for Optimization of Clinical Trial Success Prediction in Pharmaceutical R&D

统计自然语言处理用于药理研发中临床试验成功预测的优化

Michael R. Doane

专题命中 领域大模型 :LLM(abstract);分类 cs.CL、cs.LG

AI总结 本文利用统计自然语言处理技术开发了基于BioBERT的模型,以优化神经科学领域临床试验成功预测,提升研发决策效率。

Comments Doctor of Engineering Praxis Dissertation, The George Washington University. 122 pages. Present affiliation: Iambic Therapeutics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04139 2025-12-02 cs.CL 61%

Are Lexicon-Based Tools Still the Gold Standard for Valence Analysis in Low-Resource Flemish?

基于词典的工具在低资源弗拉芒语情感分析中仍然处于黄金标准吗?

Ratna Kandala, Katie Hoemann

专题命中 领域大模型 :LLM(abstract,journal_ref);分类 cs.CL

AI总结 本研究探讨了基于词典的工具在低资源弗拉芒语情感分析中的有效性,发现尽管LLM有所进步,但目前仍无法准确捕捉自发叙述中的情感价值,强调了开发文化适应性模型的必要性。

Journal ref Presented at LLM Evaluation Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00048 2025-12-02 cs.RO cs.AI 57%

Causal Reinforcement Learning based Agent-Patient Interaction with Clinical Domain Knowledge

基于因果强化学习的Agent-患者交互与临床领域知识

Wenzheng Zhao, Ran Zhang, Ruth Palan Lopez, Shu-Fen Wung, Fengpei Yuan

专题命中 领域大模型 :LLM(abstract);分类 cs.AI

AI总结 本文提出基于因果强化学习的Agent-患者交互框架,通过整合因果发现与推理提升医疗场景中的决策可解释性和适应性。

Comments Accepted by AAAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07727 2025-12-02 cs.CY 50%

Artificial Intelligence Tools Expand Scientists' Impact but Contract Science's Focus (Just accepted by Nature, to be online soon)

人工智能工具扩大科学家的影响但缩小科学的焦点(刚被《自然》接受,即将上线)

Qianyue Hao, Fengli Xu, Yong Li, James Evans

专题命中 领域大模型 :language model(abstract)

AI总结 人工智能工具提升了科学家的个人影响力,但限制了科学整体的探索范围,呈现出个人与集体利益之间的矛盾。

详情

展开后加载摘要…

URL PDF HTML 收藏