arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 925 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 925 篇

2608.06842 2026-08-10 econ.GN q-fin.EC 新提交 78%

Tabular Foundation Models and the Unity of Economic Behaviour

表格型基础模型与经济行为的统一性

Victor H. Aguiar

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 本研究通过统一选择实验,用冻结的表格型基础模型恢复决策者隐藏选择,再估计随机效用模型,构建出适用于所有经济行为领域的统一模型,提升了选择预测效果。

Comments 56 pages, 4 figures, and 19 tables, including the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05763 2026-08-07 physics.geo-ph 新提交 78%

Foundation Model-Assisted Full Waveform Inversion

基础模型辅助的全波形反演

Mustafa Alfarhan, Matteo Ravasi, Fuqiang Chen, George Turkiyyah, David Keyes

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 该研究提出用预训练地震基础模型SeisLM的特征构建全波形反演早期目标函数,通过实验验证其可避免周期跳变,提升反演效果,优于传统及混合损失工作流。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28771 2026-08-03 cs.CV 新提交 78%

Do Medical Foundation Models Generalize on the African Brain?

医学基础模型能否在非洲脑部数据上实现泛化?

Kaouther Mouheb, Gonzalo Esteban Mosquera Rojas, Juancito van Leeuwen, Stefan Klein, Esther E. Bron

机构 * Erasmus MC(伊拉斯姆斯大学医学中心)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 该研究评估医学基础模型在非洲脑部MRI数据上的泛化性,发现其无固有偏见,性能差异多源于数据集规模,核心是非洲神经影像数据集不足。

Comments Submitted to the AFRICAI workshop (Held in conjunction with MICCAI 2026, Strasbourg, France)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26276 2026-07-30 cs.CV physics.med-ph 新提交 78%

Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

比较基于基础模型的嵌入与传统方法在头颈部癌远处转移预测中的性能

Erich Schmitz, Meixu Chen, Bowen Jing, Jing Wang

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 本研究以RADCURE数据集2327例HNC患者术前CT图像为基础,对比CT基础模型嵌入、放射组学等特征集预测远处转移的性能,发现CT基础模型嵌入性能更优,可作为传统放射组学的替代方案

Comments 27 pages including supplemental materials, 5 main figures, 2 supplemental figures, 5 main tables, 7 supplemental tables. Poster Abstract at 2026 AAPM Meeting and Exhibition

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25651 2026-07-29 cs.PL cs.SE 新提交 78%

Demystifying Deep Learning Compiler Frontend Bugs: An LLM-Aided Empirical Study

揭开深度学习编译器前端错误的神秘面纱:一项基于大语言模型辅助的实证研究

Xinyi Yuan, Wei Chen, Jinyi Liu, Pengyu Chen, Jun Wei, Guoquan Wu, Jiaxin Zhu, Tao Huang

专题命中 领域大模型 :LLM(title,abstract)

AI总结 对PyTorch 2默认深度学习编译器前端TorchDynamo的fBug展开首次系统实证研究,借助领域知识增强的大语言模型辅助方法,分析fBug并构建分类法,生成测试用例,发现多个新fBug,为深度学习编译器开发和测试提供见解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25603 2026-07-29 cs.SE 新提交 78%

Input Relation Prompting for Metamorphic Testing on Query-Based Systems

基于输入关系提示的查询系统变质测试

Eng-Shen Tu, Shin-Jie Lee

专题命中 领域大模型 :prompting(title,abstract)

AI总结 针对查询系统测试难题,提出通过提示输入输出关系的变质关系识别方法,不依赖预定义测试用例或基本事实,可结合其他测试方法,经案例研究验证其适用性与潜力,推动变质测试发展,提高测试效率。

Comments 18 pages

Journal ref Journal of Information Science and Engineering, Vol. 41, No. 1, pp. 43-60 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23718 2026-07-28 cs.IR 新提交 78%

Melo: A Production LLM-Powered Music Recommendation Agent

Melo:一个由大语言模型驱动的音乐推荐代理

Shijia Wang, Da Guo, Qiang Xiao, Fanghui Bi, Weisheng Li, Dongjing Wang, Chuanjiang Luo

专题命中 领域大模型 :LLM(title,abstract)

AI总结 研究基于网易云音乐部署的大语言模型驱动音乐推荐代理Melo,针对实体幻觉和长尾退化故障模式,采用推理时实体基础和反思性重试机制,经测试在播放列表留存和参与度指标上有提升,强调运行时机制对推荐进展的重要性。

Journal ref Proceedings of the 20th ACM Conference on Recommender Systems (RecSys '26), September 27-October 02, 2026, Minneapolis, MN, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13386 2026-07-16 cs.CV 新提交 78%

FM$^2$: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging

FM$^2$:用于异构多模态医学成像的统一联邦基础模型

Shengchao Chen, Ting Shu

机构 * School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Australian AI Institute, University of Technology Sydney(悉尼科技大学澳大利亚人工智能研究所)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 针对医学成像基础模型构建中隐私与任务统一问题,提出FM$^2$框架,通过从头训练核心主干、结合预训练编码器、配备双混合专家模块及正则化器,并引入字幕增强学习,实现跨模态泛化,优于现有联邦基线。

Comments Accepted by ACM MM 2026 (Main Track): the 34th ACM International Conference on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12340 2026-07-15 cs.SE cs.CR 新提交 78%

Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents

不存在的技能:对大语言模型代理中幻觉技能推荐的大规模研究

Weifeng Yuan, Wenbo Guo, Feng Dong, Haoyu Wang, Yang Liu

专题命中 领域大模型 :LLM(title,abstract)

AI总结 研究大语言模型代理中技能名称幻觉漏洞,通过大规模测量发现各配置均有此问题,系统生成众多幻觉名称且非随机,测试的模型级防御存在安全与可用性冲突,修复需全生态系统结构改变。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02300 2026-07-10 cs.CV cs.SE 新提交 78%

Search-based Testing of Vision Language Models for In-Car Scene Understanding

基于搜索的车内场景理解视觉语言模型测试

Lev Sorokin, Chen Yang, Ken E. Friedl, Andrea Stocco

机构 * BMW Group, Technical University of Munich(宝马集团、慕尼黑技术大学) Technical University of Munich(慕尼黑技术大学) Technical University of Munich, fortiss GmbH(慕尼黑技术大学、fortiss GmbH)

专题命中 领域大模型 :language model(title,abstract)

AI总结 提出ISU-Test方法,结合渲染场景生成与搜索测试,通过优化场景参数自动生成多样化车内场景,评估VLM在问答和字幕任务中的性能,相比随机生成故障率提高10倍,故障覆盖率提高3.6倍。

Comments Accepted at the Industry Track of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04199 2026-07-07 cs.CV 新提交 78%

Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

用于3D医学视觉基础模型的拓扑驱动可迁移性估计

Jiaqi Tang, Shaoyang Zhang, Fandong Zhang, Shu Zhang, Yang Liu, Qingchao Chen

机构 * National Institute of Health Data Science, Peking University(北京大学健康数据科学研究所) Institute of Medical Technology, Peking University(北京大学医学技术研究所) Deepwise Co., Ltd.(深度智医科技有限公司) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 针对医学视觉基础模型选择难题,提出非参数、拓扑驱动框架,通过最小生成树从密集特征与语义标签的稀疏1-骨架图对齐中估计可迁移性,包含局部和全局互补尺度,提升评估效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03810 2026-07-07 cs.CV 新提交 78%

TestMate: Test-Time Domain Adaptation Aided by Lightweight Vision Foundation Model

TestMate:由轻量级视觉基础模型辅助的测试时领域适应

Dimitrios Fotiou, Vasileios Mygdalis, Ioannis Pitas

机构 * Aristotle University of Thessaloniki(塞萨洛尼基亚里士多德大学)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 研究测试时领域适应问题,提出TestMate框架,利用轻量级视觉基础模型泛化能力,通过无参数竞争融合方案实时适应,克服现有方法局限,可独立或集成提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03647 2026-07-07 cs.CV 新提交 78%

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

医学视觉语言模型真的能“看”吗?用于视觉依赖型医学视觉语言模型的反事实基础框架和硬负对比训练

Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist, Jia Wu

机构 * The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) Eisai Inc.(卫材株式会社) IISc, Bangalore(印度科学研究所班加罗尔分校) Cohere Labs Community(Cohere实验室社区)

专题命中 领域大模型 :language model(title,abstract)

AI总结 探讨医学视觉语言模型是依据视觉证据推理还是利用文本捷径,引入反事实评估框架和对比检索增强学习方法,提升模型视觉依赖能力并揭示跨域诊断差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03194 2026-07-07 cs.SE 新提交 78%

TATG: Tracking-Aware Testing Objective for LLM-based Test Generation

TATG:基于大语言模型的测试生成的跟踪感知测试目标

Guancheng Wang, Qinghua Xu, Lionel C. Briand

专题命中 领域大模型 :LLM(title,abstract)

AI总结 针对复杂Java方法自动化单元测试生成难的问题,TATG提出基于大语言模型的单元测试生成方法,引入统一目标表示跟踪测试需求,采用两阶段工作流程,实验证明其相比其他方法有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20723 2026-07-07 cs.CV 新提交 78%

Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images

医学视觉语言模型 HuluMed 和 MedGemma 以及通用聊天机器人 Gemma 3、ChatGPT Plus 和 Claude Pro 在真实未见伤口图像上的评估

Yunzhe Xue, Mohammed Saim Ahmed Quadri, Neal Panse, Justin W. Ady, Usman Roshan

机构 * Department of Computer Science, New Jersey Institute of Technology(新泽西理工学院计算机科学系) Vascular and Endovascular Surgery, Robert Wood Johnson Hospital(罗伯特·伍德·约翰逊医院血管外科) Department of Data Science, New Jersey Institute of Technology(新泽西理工学院数据科学系)

专题命中 领域大模型 :language model(title,abstract)

AI总结 本研究评估了六种视觉语言模型在慢性伤口分析任务上的表现,发现通用模型 ChatGPT 和 Claude 显著优于医学专用模型,表明广泛的多模态推理能力比领域知识更重要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01049 2026-07-02 cs.CV 新提交 78%

GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models

GenAU: 基于语言的工业异常理解与视觉语言模型

Hongkuan Zhou, Tristan Rehm, Nadeem Nazer, Lavdim Halilaj, Jingcheng Wu, Steffen Staab

机构 * Corporate Research, Robert Bosch GmbH(罗伯特·博世有限公司企业研究部) Otto-von-Guericke-University Magdeburg(马格德堡奥托·冯·格里克大学) University of Southampton(南安普顿大学)

专题命中 领域大模型 :language model(title,abstract)

AI总结 提出GenAU框架,统一图像级检测、像素级分割、多类型异常检测和缺陷分析,通过两个分割标记实现语言引导的定位,在VisA和Real-IAD上取得最优零样本检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08420 2026-06-26 cs.CV 新提交 78%

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

CheXanatomy: 面向胸部X光片的解剖感知视觉-语言建模

Sergios Gatidis, Curtis Langlotz, Christian Bluethgen

机构 * Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University(斯坦福大学医学与影像人工智能中心) Department of Radiology, Stanford University(斯坦福大学放射学系)

专题命中 领域大模型 :language model(title,abstract)

AI总结 提出CheXanatomy框架,通过自回归令牌空间监督将解剖知识融入预训练视觉-语言模型,实现解剖分割,在合成和真实X光片上性能媲美U-Net,并提升域迁移鲁棒性和样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21605 2026-06-23 cs.CV 新提交 78%

$μ$Match: Foundation Models for Semi-supervised Learning and Domain Adaptation in EM

$\mu$Match:电子显微镜中半监督学习和领域适应的基础模型

Marei Freitag, Olesia Korchevaia, Luca Freckmann, Anwai Archit, Constantin Pape

机构 * Life and Medical Sciences Institute (LIMES), University of Bonn, Germany(波恩大学生命与医学科学研究所(LIMES)) Institute of Computer Science, Georg-August-University Göttingen, Germany(哥廷根大学计算机科学研究所)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 提出μMatch框架,利用基础模型(SAM、SAM2、μSAM、DINOv2/v3)和师生方法,在电子显微镜分割任务(线粒体、细胞核、神经突)中实现半监督学习和领域适应,显著减少标注需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18123 2026-06-23 cs.CV 新提交 78%

Predicting Immune Biomarkers with MultiModal Mixture-of-Expert Pathology Foundation Models Empowers Precision Oncology

使用多模态混合专家病理基础模型预测免疫生物标志物,赋能精准肿瘤学

Tianyu Liu, Ziqing Wang, Zhaokang Liang, Tong Ding, Peter Humphrey, Lorraine Colón-Cartagena, Emily Ling-Lin Pai, Kenneth Tou En Chang, Mohamed Kahila, Jonathan Chong Kai Liew, Tinglin Huang, Rex Ying, Kaize Ding, Faisal Mahmood, Wengong Jin

机构 * Program of Computational Biology and Bioinforamtics, Yale University(耶鲁大学计算生物学与生物信息学项目) Broad Institute of MIT and Harvard(麻省理工学院与哈佛大学博德研究所) Department of Statistics and Data Science, Northwestern University(西北大学统计与数据科学系) Department of Computer Science, Northeastern University(东北大学计算机科学系) Department of Computer Science, Harvard University(哈佛大学计算机科学系) Department of Pathology, Yale University(耶鲁大学病理学系) Department of Anatomic Pathology and Laboratory Medicine, Hospital of the University of Pennsylvania(宾夕法尼亚大学医院解剖病理学与检验医学系) Department of Pathology and Laboratory Medicine, University of California, San Francisco(加州大学旧金山分校病理学与检验医学系) Department of Pathology and Laboratory Medicine, KK Women’s and Children’s Hospital(竹脚妇幼医院病理学与检验医学系) Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩雷尔曼医学院生物统计学、流行病学与信息学系)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 提出MixTIME多模态基础模型,采用混合专家架构整合不同模态的病理基础模型,从HE全切片图像预测多重免疫荧光蛋白表达,在17个蛋白标记物上达到最优性能,并增强空间域识别、生存预测等下游任务。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18749 2026-06-18 cs.CV 新提交 78%

Toward Training-Free Zero-Shot Anomaly Detection in 3D Medical Images: A Batch-Based Approach Using 2D Foundation Models

迈向3D医学图像的无训练零样本异常检测:基于批次的方法使用2D基础模型

Tai Le-Gia

机构 * Chungnam National University(忠南大学)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 提出CS3F框架,利用2D基础模型对3D医学图像进行零样本异常检测,通过沿多轴分解、切片编码和跨主体相似性计算异常分数,并引入粗到细的分词策略减少信号衰减。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17436 2026-06-17 cs.CV 新提交 78%

UoU: A Universal Fingerprint Foundation Model Based on Large-Scale Unsupervised Learning

UoU:基于大规模无监督学习的通用指纹基础模型

Xiongjun Guan, Jianjiang Feng, Jie Zhou

机构 * Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 提出UoU指纹基础模型,通过多级表示层次和结合监督、弱监督与无监督的训练策略,实现跨传感器、质量和应用的通用特征提取。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14068 2026-08-17 cs.IR cs.AI 新提交 77%

MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation

MACS:面向可靠会话式电商推荐的混合多智能体框架

Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi

专题命中 领域大模型 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对固定目录场景下会话式电商推荐的可靠性问题,提出混合多智能体框架MACS,其在单轮、多轮基准测试中均展现出更强的约束合规性与推荐性能。

Comments 9 pages, 2 figures, 8 tables. Will be presenting at Stanford Trust&Safety Conference, already presented at Stanford Market AI Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12671 2026-08-14 cs.AI cs.CC 新提交 77%

On the Expressive Power of Transformers

Transformer的表达能力研究

Phokion Kolaitis, Rik Sengupta

机构 * University of California Santa Cruz(加利福尼亚大学圣克鲁兹分校) IBM Research(IBM研究院)

专题命中 领域大模型 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文基于电路复杂性的概念与方法,概述了界定Transformer表达能力的部分精选研究结果,助力校准其作为语言识别器的表达能力。

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11259 2026-08-13 cs.CY cs.AI 新提交 77%

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

提升K-12教育中AI辅导质量的方法

Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 本文针对K-12教育中AI辅导质量提升问题,以Khanmigo为研究对象,介绍了相关衡量指标与实验,阐述了模型、提示工程等方面对指标产生积极影响的改动。

Comments 15 pages. Accepted at AIED 2026 (27th International Conference on Artificial Intelligence in Education). Published version: Artificial Intelligence in Education, LNCS vol. 16582, Springer, Cham, first online 25 June 2026 (cite as 2027)

Journal ref In: Blanchard, E.G., Chen, G., Chi, M., Isotani, S. (eds) Artificial Intelligence in Education. AIED 2026. Lecture Notes in Computer Science, vol 16582. Springer, Cham (2027)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07040 2026-08-10 cs.AI 新提交 77%

Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling

并非所有问题都最适合建模为MILP:以DSL为核心的灵活且精准的优化建模框架

Shaofeng Zhang, Hongyuan Su, Qingwen Peng, Zefang Zong, Shengcai Liu, Ke Tang, Yong Li

专题命中 领域大模型 :LLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 针对现有优化建模框架过度依赖MILP的缺陷,提出以DSL为核心的OptiDSL框架,通过LLM实现自然语言到标准化DSL的映射,在44种COP类型基准测试中显著优于MILP框架,建模准确率和效率大幅提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04366 2026-08-06 cs.CR cs.AI 新提交 77%

Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework

对抗智能体系统中的知识篡改:一种拜占庭容错的安全协同RAG框架

Zhaoqi Wang, Daqing He, Zijian Zhang, Ye Liu, Jiamou Liu, Zhirui Zeng, Zhan Qin, Zhen Li, Xin Li, Hongwei Yao, Jincheng An, Yong Liu, Yi Li, Qi Sun, Xiulei Liu, Liehuang Zhu

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对RAG系统面临的知识篡改攻击问题,提出拜占庭容错的安全协同RAG框架SecureCollaRAG,通过多源知识验证机制与动态GNN可信度评分实现攻击防护,在非IID数据下保持鲁棒性。

Journal ref Proceedings of the ACM Web Conference 2026, pages 2661-2672, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01711 2026-08-04 cs.AI 新提交 77%

Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness

使用智能体管控系统构建用于元分析综合的可执行分析知识表示

Lingbo Li, Anuradha Mathrani, Teo Susnjak

机构 * School of Mathematical and Computational Sciences(数学与计算科学学院) Massey University(梅西大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究提出EAKR,在智能体管控系统MetaSynDec中实施后,可高效构建元分析所需的可执行分析知识表示,其性能优于直接大型语言模型生成,验证了相关方法的可行性与优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21887 2026-07-27 cs.HC cs.CL 新提交 77%

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

使用水平适配的具身对话代理降低外语焦虑

Krishan Rajaratnam, Wenbin Gan, Yuan Sun

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究针对外语焦虑影响二语习得问题,提出基于欧洲共同语言参考标准的多智能体具身对话系统,通过“生成-评估-再生”循环适配用户水平。小样本试点研究表明该系统生成的对话句子更适配学习者,虽未显著降低焦虑,但提供了相关见解。

Comments 8 pages, 6 figures, published in the proceedings of EDULEARN26

Journal ref K. Rajaratnam, W. Gan, Y. Sun (2026) TOWARDS REDUCING FOREIGN LANGUAGE ANXIETY USING LEVEL-APPROPRIATE EMBODIED CONVERSATIONAL AGENTS, EDULEARN26 Proceedings, Article 1459

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17902 2026-07-21 cs.DL cs.CL 新提交 77%

Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field

用于生物医学领域研究主题本体生成的资源高效语言模型基准测试

Tanay Aggarwal, Angelo Salatino, Francesco Osborne, Enrico Motta

机构 * Knowledge Media Institute, The Open University, Milton Keynes, UK(开放大学知识媒体研究所) Department of Business and Law, University of Milano-Bicocca, Milan, IT(米兰-比科卡大学商业与法律系)

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 本文评估五个小型开源LLMs识别生物医学概念语义关系的性能,引入MeSH-Rel-4K数据集并分析三种策略,发现针对性微调使平均F1分数显著提高,突破推理瓶颈,为构建生物医学本体提供准确自动化方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11074 2026-07-14 cs.CL 新提交 77%

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

ResearchQA:科学论文中基于引用的问答基准测试

Saba Imran, Debanjum Singh Solanky

机构 * Khoj Inc.(Khoj公司)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 该研究引入ResearchQA基准测试,含多领域多类型问答对,用于科学论文基于引用的问答评估。通过特定方法评估八个模型,发现基于引用指标区分度更高,开放权重模型接近最佳封闭模型准确率且延迟更低,还发布了相关资源。

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏