arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-23 至 2026-01-23 共收录 24 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 24 篇

2406.09843 2026-01-23 cs.SE 89%

A Comprehensive Study on Large Language Models for Mutation Testing

对大型语言模型在变异测试中的全面研究

Bo Wang, Mingda Chen, Ming Deng, Youfang Lin, Mark Harman, Mike Papadakis, Jie M. Zhang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文研究了大型语言模型在变异测试中的性能,发现其在故障检测率上有显著提升,但同时也带来了非可编译性等指标的下降。

Comments 39 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15319 2026-01-23 q-bio.NC cs.AI 88%

Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles

大语言模型作为神经多样性成人心理测量剖面的模拟代理

Francesco Chiappone, Davide Marocco, Nicola Milano

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本研究探讨大语言模型能否基于结构化访谈生成接近真实个体的心理测量反应,发现其在神经发育特征模拟中表现优异,但存在特定局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15645 2026-01-23 cs.CL 86%

Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation

迈向可靠的医疗大语言模型:在真实医疗咨询中评估和增强大语言模型信心估计的基准测试

Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学) Wuhan University(武汉大学)

专题命中 领域大模型 :language model(title,abstract);large language model(title);分类 cs.CL

AI总结 本文提出MedConf框架,通过证据引导的语言自评估方法,在医疗咨询中提升大语言模型的信心估计精度与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15306 2026-01-23 cs.AI 85%

Uncovering Latent Bias in LLM-Based Emergency Department Triage Through Proxy Variables

通过代理变量揭示基于LLM的急诊科分诊中的潜在偏见

Ethan Zhang

机构 * Palo Alto High School(帕洛阿尔托高中)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究通过代理变量揭示基于LLM的急诊科分诊中的潜在偏见,发现LLMs在输入上下文中会修改对患者严重程度的感知,表明AI系统仍需改进以确保临床应用的安全性。

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12471 2026-01-23 cs.CL cs.AI 82%

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

知何时退避:医疗大语言模型在临床不确定性中的表现

Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, Zonghai Yao

机构 * Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(马萨诸塞大学阿姆赫斯特曼宁信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(医疗组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(米纳尔计算机与信息科学学院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出MedAbstain基准,探讨医疗LLM在临床不确定性中的退避能力,发现显式退避选项能显著提升安全性,而模型规模和提示方法效果有限。

Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25856 2026-01-23 cs.CV 82%

PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection

PatchEAD: 统一工业视觉提示框架用于补丁专属异常检测

Po-Han Huang, Jeng-Lin Li, Po-Hsuan Huang, Ming-Ching Chang, Wei-Chao Chen

机构 * Inventec Corporation(Inventec公司) University at Albany, State University of New York(纽约州立大学阿尔巴尼分校)

专题命中 领域大模型 :prompting(title,abstract);foundation model(abstract)

AI总结 PatchEAD提出统一的补丁聚焦框架,实现无需训练的工业异常检测,兼容多种基础模型并提升补丁相似性鲁棒性。

Comments 10 pages, 5 figures. WACV 2026 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16108 2026-01-23 cs.AI 79%

Multimodal Climate Disinformation Detection: Integrating Vision-Language Models with External Knowledge Sources

多模态气候虚假信息检测:整合视觉-语言模型与外部知识源

Marzieh Adeli Shamsabad, Hamed Ghodrati

机构 * CRIM

专题命中 领域大模型 :language model(title,abstract);分类 cs.AI

AI总结 本文提出整合视觉-语言模型与外部知识源,以提升多模态气候虚假信息检测的准确性与实时性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15721 2026-01-23 cs.IR cs.AI 77%

CoNRec: Context-Discerning Negative Recommendation with LLMs

CoNRec:基于大语言模型的上下文辨识负向推荐

Xinda Chen, Jiawei Wu, Yishuang Liu, Jialin Zhu, Shuwen Xiao, Junjun Zheng, Xiangheng Kong, Yuning Jiang

机构 * Alibaba Inc.(阿里巴巴公司) Fudan University(复旦大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 CoNRec提出了一种基于大语言模型的负向反馈建模框架,通过上下文辨识模块和渐进式训练方法,提升对用户负向偏好的建模能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15558 2026-01-23 cs.CL 77%

From Generation to Collaboration: Using LLMs to Edit for Empathy in Healthcare

从生成到协作:利用LLMs编辑以增强医疗中的同理心

Man Luo, Bahareh Harandizadeh, Amara Tariq, Halim Abbas, Umar Ghaffar, Christopher J Warren, Segun O. Kolade, Haidar M. Abdul-Muhsin

机构 * Science team(科学团队) Abridge Mayo Clinic(梅奥诊所) Department of AI & Informatics(人工智能与信息学部门) Urology Department(泌尿科部门)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究利用LLMs作为编辑工具,通过优化医生书面回应增强医疗中的同理心,同时保持事实准确性,提出新的评估指标并验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15339 2026-01-23 cs.SE cs.AI 77%

Lost in Transcription: How Speech-to-Text Errors Derail Code Understanding

迷失在转录中:语音转文本错误如何阻碍代码理解

Jayant Havare, Ashish Mittal, Srikanth Tamilselvam, Ganesh Ramakrishnan

机构 * IIT Bombay(印度理工学院班加罗尔分校) IBM Research(IBM研究)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一个多语言语音驱动的代码理解框架,通过改进语音转录和代码理解,解决多语言和语音驱动编程工具中的转录错误问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20097 2026-01-23 q-fin.CP cs.AI 77%

Can LLMs Identify Tax Abuse?

大语言模型能识别税务滥用吗?

Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究大语言模型能否识别和分析美国税务最小化策略,发现其能生成新颖策略,可能革新税务机构应对税务滥用的方式。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05793 2026-01-23 cs.CY cs.AI cs.CL 73%

MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education

MedSimAI:通过模拟和形成性反馈提升医学教育中的刻意练习

Yann Hicke, Jadon Geathers, Kellen Vu, Justin Sewell, Claire Cardie, Jaideep Talwalkar, Dennis Shung, Anyanate Gwendolyne Jack, Susannah Cornes, Mackenzi Preston, Rene Kizilcec

机构 * Cornell University(康奈尔大学) UCSF School of Medicine(旧金山加利福尼亚大学医学院) Yale School of Medicine(耶鲁大学医学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 MedSimAI通过模拟和形成性反馈提升医学教育中的刻意练习,通过AI生成临床互动并提供自动评估,提高病史采集和沟通技能。

Comments Accepted to LAK 2026; 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16018 2026-01-23 cs.CL 70%

Mecellem Models: Turkish Models Trained from Scratch and Continually Pre-trained for the Legal Domain

Mecellem模型:从零开始训练并持续预训练的土耳其法律领域模型

Özgür Uğur, Mahmut Göksu, Mahmut Çimen, Musa Yılmaz, Esra Şavirdi, Alp Talha Demir, Rumeysa Güllüce, İclal Çetin, Ömer Can Sağbaş

专题命中 领域大模型 :language model(abstract);post-training(abstract);分类 cs.CL

AI总结 Mecellem模型通过从零训练和持续预训练,实现了在土耳其法律领域中的高效领域适应,取得前三名成绩并提升生产效率。

Comments 16 png, 1 tex, 1 bib

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15476 2026-01-23 cs.AI cs.PF 70%

Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases

可靠性设计:量化并消除大语言模型中的制造风险。从生成式到咨询式AI:法律领域中的比较分析及对高风险知识库的启示

Alex Dantart

机构 * Humanizing Internet

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出通过减少幻觉提升大语言模型在法律领域的可靠性,引入两种可靠性指标,并展示先进的 RAG 系统有效降低伪造事实率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15429 2026-01-23 cs.CL 70%

Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs

在增强医疗LLM中的领域特定知识图谱

Sydney Anuyah, Mehedi Mahmud Kaushik, Hao Dai, Rakesh Shiradkar, Arjan Durresi, Sunandan Chakraborty

机构 * Luddy School of Informatics, Computing, and Engineering, Indiana University, Indianapolis, IN, USA(信息学、计算与工程学院,印第安纳大学,印第安纳波利斯,IN,USA) School of Medicine, Indiana University, Indianapolis, IN, USA(医学学院,印第安纳大学,印第安纳波利斯,IN,USA) Department of Biomedical Engineering and Informatics, Indiana University, Indianapolis, IN, USA(生物医学工程与信息学系,印第安纳大学,印第安纳波利斯,IN,USA)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 该研究探讨了在医疗LLM中利用领域特定知识图谱提升检索增强生成的效果,发现精准匹配的图谱检索优于随意联合,且模型大小和温度对性能影响各异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15663 2026-01-23 cs.CR cs.AI cs.LG 62%

TempoNet: Learning Realistic Communication and Timing Patterns for Network Traffic Simulation

TempoNet: 学习真实通信和时间模式以进行网络流量模拟

Kristen Moore, Diksha Goel, Cody James Christopher, Zhen Wang, Minjune Kim, Ahmed Ibrahim, Ahmad Mohsin, Seyit Camtepe

专题命中 领域大模型 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 TempoNet通过结合多任务学习和多标记时序点过程,生成高保真的网络流量模拟,有效解决真实网络时序和通信动态建模的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15506 2026-01-23 cs.CL cs.LG 62%

ViT Registers and Fractal ViT

ViT 注册与分形 ViT

Jason Chuan-Chih Chou, Abhinav Kumar, Shivank Garg

机构 * Cohere Labs(Cohere实验室) Mathematics and Computing(数学与计算) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) Artificial Intelligence and Data Science(人工智能与数据科学)

专题命中 领域大模型 :language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出分形 ViT,通过引入注意力掩码打破标记对称性,探讨注册机制对 ViT 性能的影响,发现其效果可能受规模、领域或应用限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15798 2026-01-23 cs.AI 57%

VitalDiagnosis: AI-Driven Ecosystem for 24/7 Vital Monitoring and Chronic Disease Management

VitalDiagnosis:基于AI的生态系统,用于24/7生命体征监测和慢性病管理

Zhikai Xue, Tianqianjin Lin, Pengwei Yan, Ruichun Wang, Yuxin Liu, Zhuoren Jiang, Xiaozhong Liu

专题命中 领域大模型 :LLM(abstract);分类 cs.AI

AI总结 VitalDiagnosis利用AI技术,通过整合可穿戴设备数据与大语言模型,实现24/7生命体征监测和慢性病管理的主动参与,提升患者自我管理能力并减少临床工作量。

Comments Accepted by AAAI 2026 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15745 2026-01-23 cs.CL 57%

Hallucination Mitigating for Medical Report Generation

缓解医疗报告生成中的幻觉

Ruoqing Zhao, Runze Xia, Piji Li

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(信息产业部模式分析与机器智能重点实验室) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 领域大模型 :language model(abstract);分类 cs.CL

AI总结 KERM框架通过知识检索、净化模块和细粒度奖励,有效缓解医疗报告生成中的幻觉问题,提升报告质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15575 2026-01-23 cs.HC cs.AI cs.IR 57%

PromptHelper: A Prompt Recommender System for Encouraging Creativity in AI Chatbot Interactions

PromptHelper: 一种促进AI聊天机器人交互中创造力的提示推荐系统

Jason Kim, Maria Teleki, James Caverlee

机构 * Texas A\&M University College Station Texas U.S.A. Texas A\&M University

专题命中 领域大模型 :prompting(abstract);分类 cs.AI

AI总结 PromptHelper通过提供语义多样的提示建议,帮助用户在AI聊天机器人交互中提升创造力和表达性,同时减少认知负担。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15308 2026-01-23 cs.HC cs.AI 57%

When Generative AI Meets Extended Reality: Enabling Scalable and Natural Interactions

当生成式AI遇见扩展现实:实现可扩展和自然的交互

Mingyu Zhu, Jiangong Chen, Bin Li

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 本文探讨生成式AI与扩展现实的结合,通过三个用例展示如何通过语言驱动交互和自动化内容生成解决XR在可扩展性和自然交互方面的挑战。

Comments Accepted by IEEE Internet Computing (Oct. 2025); published in IEEE Xplore (Jan. 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22777 2026-01-23 cs.CL 57%

MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators

MEDAL:一个用于评估LLM作为多语言开放领域对话评估器的框架

John Mendonça, Alon Lavie, Isabel Trancoso

专题命中 领域大模型 :LLM(abstract);分类 cs.CL

AI总结 MEDAL框架通过多语言多代理方法,评估LLM作为开放领域对话评估者的性能,揭示了现有评判者在检测细微问题上的不足。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16060 2026-01-23 cs.CV 50%

ProGiDiff: Prompt-Guided Diffusion-Based Medical Image Segmentation

ProGiDiff: 基于提示引导的扩散模型医学图像分割

Yuan Lin, Murong Xu, Marc Hölle, Chinmay Prabhakar, Andreas Maier, Vasileios Belagiannis, Bjoern Menze, Suprosanna Shit

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(埃朗根-纽伦堡弗里德里希-亚历山大大学) University of Zurich(苏黎世大学)

专题命中 领域大模型 :prompting(abstract)

AI总结 ProGiDiff通过提示引导的扩散模型实现医学图像分割,结合自定义编码器和多类条件化机制,提升分割性能和跨模态适应能力。

Comments 5 pages, 4 figures. It has been accepted by IEEE ISBI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24106 2026-01-23 cs.CE 50%

UniField: Joint Multi-Domain Training for Universal Surface Pressure Modeling

UniField:面向通用表面压力建模的多领域联合训练

Junhong Zou, Zhenxu Sun, Yueqing Wang, Wei Qiu, Zhaoxiang Zhang, Xiangyu Zhu, Zhen Lei

专题命中 领域大模型 :foundation model(abstract)

AI总结 UniField通过多领域联合训练提升通用表面压力建模的准确性,尤其在数据稀缺领域表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏