arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 461 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 461 篇

2607.00011 2026-07-15 cs.IR cs.AI cs.SE 版本更新 90%

SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents

SkillSelect-Serve:面向小型LLM代理的预算可控且QoS感知的技能服务推荐与组合

Jingyuan Zheng, Dongjing Wang, Xin Zhang, Hao Chen, Youhuizi Li, Xudong Shen, Haiping Zhang, Butian Huang, Dongjin Yu, Guandong Xu

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出SkillSelect-Serve框架,将技能选择建模为服务推荐与组合问题,通过双粒度效用建模和预算约束优化,在35,353个技能和586个任务查询上优于固定top-k检索基线。

Comments 18 pages (14-page main text + appendices), 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15245 2026-07-09 cs.CL cs.HC 版本更新 90%

Practicing with Language Models Cultivates Human Empathic Communication

与语言模型练习培养人类共情沟通

Aakriti Kumar, Nalin Poungpeth, Diyi Yang, Bruce Lambert, Matthew Groh

机构 * Kellogg School of Management, Northwestern University(西北大学凯洛格管理学院) Northwestern Institute for Complex Systems, Northwestern University(西北大学复杂系统研究所) Ryan Institute on Complexity, Northwestern University(西北大学复杂性研究院) Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Department of Communication Studies, Northwestern University(西北大学传播学系) Department of Computer Science, Northwestern University(西北大学计算机科学系)

专题命中 领域大模型 :LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究通过实验平台探讨共情沟通技能的提升,发现简短的LLM指导干预能有效提升参与者与规范共情沟通模式的匹配度,并发现沉默共情效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10574 2026-06-16 cs.AI 版本更新 90%

LLM Jaggedness Unlocks Scientific Creativity

LLM Jaggedness Unlocks Scientific Creativity

Shray Mathur, J. Anibal Boscoboinik, Esther H. R. Tsai, Kevin G. Yager

机构 * Center for Functional Nanomaterials, Brookhaven National Laboratory(功能纳米材料中心,布鲁赫萨尔国家实验室)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究了大型语言模型(LLMs)在科学创意生成中的不规则进步现象,提出了SciAidanBench基准测试集,通过评估不同模型在科学问题上的创意生成能力,揭示了模型在跨任务、提示和领域层面的不均衡表现,并展示了如何通过推理计算、知识聚合和头脑风暴等机制利用这种不规则性来提升科学创造力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10120 2026-06-11 cs.IR cs.AI cs.HC 版本更新 90%

MetaPlate: Counterfactual-Guided RAG-LLM Tool for Personalized Food Recommendation and Hyperglycemia Prevention

MetaPlate: 反事实引导的RAG-LLM工具用于个性化食物推荐和高血糖预防

Asiful Arefeen, Carol Johnston, Hassan Ghasemzadeh

机构 * College of Health Solutions, Arizona State University(亚利桑那州立大学健康解决方案学院) School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.AI

AI总结 提出MetaPlate框架,结合反事实解释、机器学习预测和RAG-LLM,生成个性化膳食建议以预防餐后高血糖,经注册营养师评估证明其可行性和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21239 2026-06-09 cs.CL 版本更新 90%

A Unified LLM-Adaptable Framework for Cold-Start Cognitive Diagnosis

面向冷启动认知诊断的统一LLM可适配框架

Zihan Yao, Chentao Song, Yu He, Tianyu Qi, Jian Zhang, Weiping Fu, Jun Liu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出LMCD框架,通过知识扩散和语义-认知融合两阶段,利用大语言模型增强冷启动场景下的认知诊断性能。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10896 2026-06-08 cs.CL 版本更新 90%

DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference

DialDefer: 检测和缓解LLM对话性遵从的框架

Parisa Rabbani, Priyam Sahoo, Ruben Mathew, Aishee Mondal, Harshita Ketharaman, Nimet Beyza Bozdag, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.CL

AI总结 提出DialDefer框架,通过对话性遵从分数检测和缓解LLM在对话评估中因提问框架导致的判断偏移,发现框架效应显著但准确率稳定,且模型对人类与AI的不同归因产生最大偏移。

Comments 10 pages main content, 7 figures, 35 pages total with appendix

Journal ref ACL 2026 - Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14732 2026-08-17 cs.LG cs.AI cs.CV eess.IV 版本更新 90%

INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT

INFORM-CT:整合LLM和VLM用于腹部CT的偶发发现管理

Idan Tankel, Nir Mazor, Rafi Brada, Christina LeBedis, Guy ben-Yosef

机构 * GE Healthcare Technology and Innovation Center(GE医疗技术与创新中心) Boston Medical Center(波士顿医疗中心)

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于LLM和VLM的计划-执行框架,用于提高腹部CT偶发发现的检测、分类和报告效率与精度,通过自动化流程提升临床应用效果。

Comments Spotlight presentation at the 9th International Conference on Medical Imaging with Deep Learning (MIDL) 2026 Additional code and implementation details available at https://idan-tankel.github.io/InformCT_ProjectPage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01826 2026-08-17 cs.CL cs.AI 版本更新 90%

Leveraging Few-Shot Learning and Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature

利用小样本学习与大语言模型从科学文献中分析不同生物性别间的血压差异

Yuting Guo, Seyedeh Somayyeh Mousavi, Reza Sameni, Abeed Sarker

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本研究利用小样本学习、LLaMA3及GPT-3.5等大语言模型,从PubMed文献中提取血压相关信息,分析不同生物性别间的血压差异,生成可视化图表开展研究。

Comments Accepted by the journal of Computers in Biology and Medicine

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09548 2026-08-12 cs.CL cs.AI cs.CY 版本更新 90%

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

ELBench:面向教育场景的大语言模型多维基准

Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Guangtao Zhai, Aimin Zhou

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);post-training(abstract);分类 cs.CL、cs.AI

AI总结 ELBench是首个评估教育大模型四项核心要求的综合基准,评估发现通用模型综合表现相近但模块优势不同,中国模型在安全性模块领先,教育专用模型未在教育模块占优,高阶培养存在系统性盲点。

Comments 13 pages, 6 figures, 8 tables. Benchmark data: https://huggingface.co/datasets/ZeroLoss-Lab/ELBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12393 2026-07-27 cs.CL cs.AI 版本更新 90%

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

MedKGent:用于构建随时间演变的医学知识图谱的大语言模型智能体框架

Duzhen Zhang, Zixiao Wang, Zhong-Zhi Li, Yahan Yu, Shuncheng Jia, Jiahua Dong, Haotian Xu, Xing Wu, Yingying Zhang, Tielin Zhang, Jie Yang, Xiuying Chen, Le Song

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德大学人工智能学院) University of Chinese Academy of Sciences(中国科学院大学) Kyoto University(京都大学) Tsinghua University(清华大学) East China Normal University(华东师范大学) Center for Excellence in Brain Science and Intelligence Technology(脑科学与智能技术卓越中心) Brigham and Women’s Hospital, Harvard Medical School(哈佛医学院布里特妇女医院) GenBio AI

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 研究针对医学文献增长带来的知识结构化挑战,引入MedKGent框架,利用PubMed摘要通过两个智能体每日增量构建医学知识图谱,经评估其三元组有效性高,能显著改进大语言模型在医学问答基准上的检索增强生成。

Comments Accepted by Npj Digital Medicine

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14892 2026-06-15 cs.LG cs.AI 版本更新 90%

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

LLM能否准确评分医学诊断和临床推理?

Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett

机构 * Wits MIND Institute, University of the Witwatersrand, Johannesburg, South Africa(维特士心理研究所,沃斯兰德大学,约翰内斯堡,南非) Grai Labs, Cape Town, South Africa(格雷实验室,开普敦,南非) South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(南非医学研究理事会疫苗和传染病分析研究组,健康科学学院,沃斯兰德大学,约翰内斯堡,南非) Department of Internal Medicine, Charlotte Maxeke Johannesburg Academic Hospital, and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(内科学系,查理·马克斯凯约翰内斯堡学术医院,以及健康科学学院,沃斯兰德大学,约翰内斯堡,南非) Department of Paediatrics and Child Health, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(儿科学与儿童健康系,健康科学学院,沃斯兰德大学,约翰内斯堡,南非) Wits MIND Institute, University of the Witwatersrand, Johannesbu(维特士心理研究所,沃斯兰德大学,约翰内斯堡)

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究使用LLM陪审团对300例低收入和中等收入国家医院病例的3334个诊断进行评分,发现校准后的LLM评分与专家评分高度一致,且严重错误风险更低,可作为可靠的评估代理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09546 2026-08-18 q-fin.TR cs.SI 版本更新 89%

A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading

一种用于指导零样本加密货币交易的反思型大语言模型智能体

Yuan Li, Bingqiao Luo, Qian Wang, Nuo Chen, Xu Liu, Bingsheng He

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract)

AI总结 本研究开发了结合链上链下数据分析与反思机制的LLM智能体CryptoTrade,拓展了LLM在加密货币交易的应用,建立了相关策略基准,其收益表现优于传统策略和时间序列基线。

Comments Published at EMNLP 2024 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30384 2026-08-14 astro-ph.IM astro-ph.GA 版本更新 89%

The EPS Research Astro-RAG Platform: A Unified Open-Science Infrastructure for Cross-Epoch Astrophysical Kinematic Analysis, LLM-Assisted Research Workflows, and Educational Outreach

EPS Research Astro-RAG平台:用于跨纪元天体物理运动学分析、LLM辅助研究工作流程和教育推广的统一开放科学基础设施

David C. Flynn

专题命中 领域大模型 :LLM(title,title_cn)

AI总结 提出一个开放科学平台,整合四个跨纪元天体物理语料库、120个可执行Jupyter笔记本示例,并利用LLM辅助检索增强生成(RAG)研究工作流程,通过ω运动学校正揭示从z=0到z~5的角速度符号反转。

Comments 5 pages, no figures, no tables. Platform release: four open astrophysical corpora (772 objects, z=0 to z~6), 120 verified Jupyter notebooks, QuickStart reproducibility pathway, and High-School Exploration Track. Zenodo DOIs: 10.5281/zenodo.19563417, 10.5281/zenodo.20320362, 10.5281/zenodo.19907765, 10.5281/zenodo.20369285. GitHub: github.com/eps-research/rag-corpus-series

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09181 2026-07-29 cs.DB 版本更新 89%

Evaluating the Practical Effectiveness of LLM-Driven Index Tuning on Microsoft SQL Server

评估基于大语言模型的索引优化的实用性效果:微软数据库优化顾问

Xiaoying Wang, Wentao Wu, Vivek Narasayya, Surajit Chaudhuri

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract)

AI总结 本文通过工业基准和真实企业工作负载评估LLM驱动的索引优化效果,对比微软数据库优化顾问(DTA),发现LLM在部分情况下能提供更优执行时间配置,但生产环境中应用仍面临性能波动、集成影响和验证成本等挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12865 2026-07-03 cs.HC 版本更新 89%

DashChat: Interactive Authoring of Performance Dashboard Design Prototypes through Conversation with LLM-Powered Agent

DashChat: 通过与LLM驱动的智能体对话交互式创作性能仪表盘设计原型

Z. Lin, S. Shen, W. Liu, C. Xin, W. Dai, S. Chen, X. Wen, X. Lan

专题命中 领域大模型 :LLM(title,title_cn)

AI总结 提出DashChat系统,通过LLM驱动的多智能体管道将文本需求转化为性能仪表盘原型,加速迭代并保证设计质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14340 2026-06-26 cs.SD 版本更新 89%

Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

通过语音-文本对齐细化伪音频提示以实现基于LLM的ASR的纯文本领域适应

Ryo Magoshi, Takashi Maekaku, Yusuke Shinohara

机构 * Kyoto University, Japan(京都大学,日本) LY Corporation, Japan(LY公司,日本)

专题命中 领域大模型 :LLM(title,title_cn)

AI总结 本文提出通过语音-文本对齐细化伪音频提示的方法,以提升基于LLM的ASR在纯文本领域适应中的性能,实验表明该方法在整体错误率和词汇覆盖率上优于现有方法。

Comments Accepted at Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12450 2026-06-09 cs.CY 版本更新 89%

Empirical Modeling of Therapist-Client Dynamics in Psychotherapy Using LLM-Based Assessments

基于LLM评估的心理治疗中治疗师-来访者动态的实证建模

Angela Chen, Siwei Jin, Canwen Wang, Holly Swartz, Tongshuang Wu, Robert E Kraut, Haiyi Zhu

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本研究利用大语言模型测量治疗师行为、关系质量和来访者结果,结合结构方程模型分析约2000小时转录数据,发现共情和探索直接促进来访者自我表露和情绪变化,而融洽关系起调节作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03997 2026-08-07 cs.CL 版本更新 89%

Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs

基于大语言模型(LLM)从全国在线评论中映射患者感知的医生特质

Junjie Luo, Rui Han, Arshana Welivita, Zeleikun Di, Jingfu Wu, Xuzhe Zhi, Ritu Agarwal, Gordon Gao

机构 * Johns Hopkins School of Medicine(约翰霍普金斯医学院) Johns Hopkins University(约翰霍普金斯大学) Carey Business School, Johns Hopkins University(约翰霍普金斯大学Carey商学院) Center for Digital Health Artificial Intelligence (CDHAI)(数字健康人工智能中心(CDHAI))

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究提出基于LLM的流程从美国百万级医生的410万条在线评论中提取10项患者感知的医生特质,揭示了全国性特质分布模式与医生原型,为相关公平性、偏见研究提供了基础。

Comments Accepted in npj Digital Medicine

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05413 2026-06-15 cs.IR cs.CL 版本更新 89%

SciDef: Datasets and Tools for Automated Definition Extraction from Scientific Literature with LLMs

SciDef:基于LLM的科学文献自动定义提取数据集与工具

Filip Kučera, Christoph Mandl, Isao Echizen, Radu Timofte, Timo Spinde

机构 * National Institute of Informatics (NII)(国立信息研究所) University of Würzburg(乌尔姆大学) University of Passau(帕萨乌大学) University of Würzburg (JMU)(乌尔姆大学)

专题命中 领域大模型 :LLM(title_cn,summary_cn);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 提出SciDef资源套件,包含人工验证的定义基准DefExtra、相似度判断DefSim及基于LLM的提取流程,通过16个语言模型评估,发现NLI匹配指标与人类判断高度一致,但相关性过滤仍是自动提取的关键瓶颈。

Comments Under Review - Submitted to CIKM 2026 Resources Track;

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17467 2026-08-04 cs.CL cs.AI cs.CY cs.HC cs.LG 版本更新 89%

Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems

理解大语言模型对写作与信息生态系统影响的计算方法

Weixin Liang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本论文以计算方法研究LLMs对写作与信息生态系统的影响,涵盖AI检测器的公平性问题、LLMs在多写作领域的采用模式及LLMs为研究者提供手稿反馈的潜力。

Comments Stanford CS PhD Dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01410 2026-06-16 cs.HC 版本更新 89%

What LLMs Must Forget to Teach Effectively: A DIY Approach to Premodern Japanese Language Pedagogy

LLM必须忘记什么才能有效教学:前现代日语教学法的DIY方法

Ariel Stilerman, Andrew Nelson, Alan Cheng, Caleb Langley, Sera Wang, Camilla Piana, Pelin Çılgın, Qianhe Qin, Teisha Nishimitsu, Liaoliao Zhang, Huiting Liu, Josh Eyre, Gavin Sherry

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract)

AI总结 本文提出一种基于大型语言模型(LLM)的DIY教学框架,通过提示工程创建定制工具,避免LLM过度解释或产生幻觉,从而促进前现代日语文学和语言课程中的主动理解与教学对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28369 2026-07-02 cs.IR cs.AI 版本更新 89%

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

基于AI基础模型的多模态与多尺度时空语义搜索与推荐

Yuanyuan Tian, Wenwen Li, Xiao Chen, Michael Brook, Michael Brubaker, Anna Liljedahl, Chitta Baral

机构 * School of Geographical Sciences and Urban Planning, Arizona State University(地理科学与城市规划学院,亚利桑那州立大学) Alaska Native Tribal Health Consortium(阿拉斯加原住民部落健康联盟) Woodwell Climate Research Center(伍德沃德气候研究中心) School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学)

专题命中 领域大模型 :foundation model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出融合大语言模型和视觉-语言模型的框架,通过CAMERA算法融合文本与视觉信息生成更丰富嵌入,以及ASTRA算法结合尺度依赖的时空相关性与语义相似性进行重排序,在环境事件文档检索中优于单模态方法。

Comments 17 pages, accepted for publication in the ACM Transactions on Spatial Algorithms and Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08636 2026-08-17 cs.CL cs.AI cs.DL cs.IR 版本更新 88%

Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach

基于大语言模型增强科学命名实体识别:一种类型驱动的多任务学习方法

Tong Bao, Yi Zhao, Heng Zhang, Chengzhi Zhang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究针对LLMs处理SciNER时因实体类型过多导致准确率低的问题,提出类型驱动的多任务学习方法TdSciNER,通过实体类型筛选、多任务学习和示例选择策略提升性能,相关方法在三个数据集上达到与全监督模型相当的效果。

Journal ref Expert Systems With Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29738 2026-08-10 cs.CL cs.AI 版本更新 88%

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

Multi-Legal-Bench: 跨司法管辖区、语言和法律传统的法律推理评估LLM

Volodymyr Ovcharov

机构 * SecondLayer

专题命中 领域大模型 :LLM(title_cn,summary_cn);pretraining(abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 提出Multi-Legal-Bench,首个跨司法管辖区法律基准,在6个国家、4个语系和1.34亿份法院判决上评估LLM,发现少样本效果跨辖区复制、无单一模型主导所有语言、跨语言迁移不遵循语言邻近性、分词器效率不显著预测跨语言准确率。

Comments 17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15176 2026-07-22 cs.AI cs.CL cs.HC 版本更新 88%

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

用于科学可视化素养的多模态大语言模型基准测试

Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

机构 * University of Notre Dame(圣母大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 研究对六个多模态大语言模型进行科学可视化素养基准测试,涵盖多种技术和任务类型。通过封闭世界协议评估闭源和开源模型,与人类参与者数据对比。发现模型表现不均,Gemini最强,开源模型低于人类基线,明确SciVis素养对评估多模态AI系统的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19021 2026-07-21 cs.CR 版本更新 88%

Trustworthy AI LLM Scalability Risk Index (LSRI): A Cybersecurity Framework Assessing Agentic-AI Security & Software Model Supply Chain Safety Boosting AI-Generated Malware Defense & Explainability Mitigating Emerging Risks of Generative AI

大语言模型在代理AI中的可扩展性风险与模型供应链安全

Kiarash Ahi, Vaibhav Agrawal, Saeed Valizadeh

专题命中 领域大模型 :LLM(title,summary_cn);RLHF(abstract)

AI总结 本文研究了大语言模型在代理AI中的可扩展性风险及模型供应链安全问题,提出LSRI指数和模型供应链框架,以提升安全关键环境下的LLM部署安全性。

Comments Accepted for publication in Journal of Computer Information Systems (2026). DOI: 10.1080/08874417.2026.2624670

Journal ref Journal of Computer Information Systems (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00260 2026-07-01 cs.CV 版本更新 88%

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

TotalFM:利用大规模常规临床放射学数据的器官分离3D-CT基础模型

Kohei Yamamoto, Tomohiro Kikuchi

机构 * Department of Radiology, Jichi Medical University(放射科,自治医科大学) Data Science Center, Jichi Medical University(数据科学中心,自治医科大学)

专题命中 领域大模型 :foundation model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出TotalFM,一种基于器官分离的3D-CT放射学基础模型,通过自动化生成器官体积与发现句子对,结合VideoMAE自监督预训练和对比学习,在零样本器官级和发现级病变分类任务中优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01693 2026-06-29 cs.DB cond-mat.mtrl-sci 版本更新 88%

LitMOF: An LLM Multi-Agent for Literature-Validated Metal-Organic Frameworks Database Correction and Expansion

LitMOF:用于文献验证的金属有机框架数据库校正与扩展的LLM多智能体系统

Honghui Kim, Dohoon Kim, Jihan Kim

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 提出LitMOF框架,利用大语言模型多智能体从原始文献验证并修复MOF数据库中的结构错误,成功修复9227个无效条目并发现8771个新MOF,通过直接空气捕获案例证明结构错误严重影响材料筛选可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26233 2026-08-06 cs.AI cs.CY 版本更新 88%

Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools

说服性与大语言模型作为法律决策工具

Oisin Suttle, David Lillis

机构 * School of Law Criminology Maynooth University Maynooth Ireland School of Computer Science University College Dublin Dublin Ireland Criminology Maynooth University School of Computer Science University College Dublin

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 研究大语言模型在法律决策中的说服性,分析其对法律论点的响应机制及影响因素,探讨模型在法律和行政领域的应用可行性。

Comments v2 is the conference version, accepted for the Proceedings of the 21st International Conference on Artificial Intelligence and Law (ICAIL 2026), (DOI: 10.1145/3836937.3837003). v3 is a substantially extended version for journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00285 2026-08-06 cs.CL 版本更新 88%

Enhancing Trustworthy Clinical Diagnosis Decision-Making in Large Language Models via Etiology-Aware Attention Supervision

通过病因感知注意力监督增强大型语言模型在临床诊断决策中的可信性

Peixian Li, Yu Tian, Ruiqi Tu, Chengkai Wu, Jingjing Ren, Jingsong Li

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 该研究提出病因感知注意力监督框架,通过引入临床病因模式对大型语言模型进行参数高效微调,在两类诊断队列上提升了诊断准确率与注意力聚焦性,增强了模型临床诊断的可信性。

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏