arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12705 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 12705 篇

2508.06732 2026-08-14 cs.HC cs.LG 77%

ClimateSOM: A Visual Analysis Workflow for Climate Ensemble Datasets

Yuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu Ma

机构 * University of California, Davis(加州大学戴维斯分校) Scripps Institution of Oceanography, University of California, San Diego(Scripps海洋研究所,加州大学圣地亚哥分校)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11259 2026-08-13 cs.CY cs.AI 新提交 77%

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

提升K-12教育中AI辅导质量的方法

Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 本文针对K-12教育中AI辅导质量提升问题,以Khanmigo为研究对象,介绍了相关衡量指标与实验,阐述了模型、提示工程等方面对指标产生积极影响的改动。

Comments 15 pages. Accepted at AIED 2026 (27th International Conference on Artificial Intelligence in Education). Published version: Artificial Intelligence in Education, LNCS vol. 16582, Springer, Cham, first online 25 June 2026 (cite as 2027)

Journal ref In: Blanchard, E.G., Chen, G., Chi, M., Isotani, S. (eds) Artificial Intelligence in Education. AIED 2026. Lecture Notes in Computer Science, vol 16582. Springer, Cham (2027)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18452 2026-08-13 cs.CL 77%

MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification

Heyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, Mark Dredze

机构 * Center for Language and Speech Processing(语言与语音处理中心) Johns Hopkins University(约翰霍普金斯大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments Added generalizability experiment and examples on non-medical free-form answer. Added ablation study for MedCorp verification corpus and MedScore decomposition prompt

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 14149-14180, San Diego, California, United States. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17629 2026-08-11 cs.HC cs.AI cs.CY 版本更新 77%

A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence

一项严格的图灵测试:评估通用人工智能的基础

Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究严格遵循图灵原始规则开展图灵测试,发现仅1名参与者误识别大型语言模型,表明其通过图灵测试的说法尚不成熟,图灵测试仍将是评估机器智能的核心手段。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07040 2026-08-10 cs.AI 新提交 77%

Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling

并非所有问题都最适合建模为MILP:以DSL为核心的灵活且精准的优化建模框架

Shaofeng Zhang, Hongyuan Su, Qingwen Peng, Zefang Zong, Shengcai Liu, Ke Tang, Yong Li

专题命中 领域大模型 :LLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 针对现有优化建模框架过度依赖MILP的缺陷,提出以DSL为核心的OptiDSL框架,通过LLM实现自然语言到标准化DSL的映射,在44种COP类型基准测试中显著优于MILP框架,建模准确率和效率大幅提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27540 2026-08-10 cs.AI 版本更新 77%

In-Context Examples Suppress Scientific Knowledge Recall in LLMs

上下文示例抑制大语言模型中的科学知识回忆

Chaemin Jang, Woojin Park, Hyeok Yun, Dongman Lee, Jihee Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Shanghai Jiao Tong University(上海交通大学)

专题命中 领域大模型 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究发现上下文示例会削弱大语言模型对科学知识的回忆能力,使模型更依赖经验模式匹配而非预训练知识。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04366 2026-08-06 cs.CR cs.AI 新提交 77%

Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework

对抗智能体系统中的知识篡改:一种拜占庭容错的安全协同RAG框架

Zhaoqi Wang, Daqing He, Zijian Zhang, Ye Liu, Jiamou Liu, Zhirui Zeng, Zhan Qin, Zhen Li, Xin Li, Hongwei Yao, Jincheng An, Yong Liu, Yi Li, Qi Sun, Xiulei Liu, Liehuang Zhu

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对RAG系统面临的知识篡改攻击问题,提出拜占庭容错的安全协同RAG框架SecureCollaRAG,通过多源知识验证机制与动态GNN可信度评分实现攻击防护,在非IID数据下保持鲁棒性。

Journal ref Proceedings of the ACM Web Conference 2026, pages 2661-2672, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01711 2026-08-04 cs.AI 新提交 77%

Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness

使用智能体管控系统构建用于元分析综合的可执行分析知识表示

Lingbo Li, Anuradha Mathrani, Teo Susnjak

机构 * School of Mathematical and Computational Sciences(数学与计算科学学院) Massey University(梅西大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究提出EAKR,在智能体管控系统MetaSynDec中实施后,可高效构建元分析所需的可执行分析知识表示,其性能优于直接大型语言模型生成,验证了相关方法的可行性与优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17480 2026-07-30 cs.AI 版本更新 77%

The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure

能力悖论:更聪明的审计员如何使多智能体系统更不安全

Qiqi Liu, Runhan Song, Shilin Ye

机构 * University of Chinese Academy of Sciences(中国科学院大学) Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所) Henan Yinzhu Safety Technology Co., Ltd.(河南亿众安全技术有限公司) Harbin Institute of Technology, Faculty of Computing(哈尔滨工业大学计算机学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 本文研究了多智能体系统中,随着工人能力的提升,系统级攻击成功率反而上升的现象,揭示了语言确定性在攻击传播中的作用,并提出异质性集成验证作为解决方案,以降低攻击成功率。

Comments 28 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07131 2026-07-30 cs.CV cs.AI 版本更新 77%

Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge

深度专家注入用于带有领域特定知识的视网膜视觉语言模型的锚定

Shuai Lu, Meng Wang, Jia Guo, Jiawei Du, Bo Liu, Shengzhu Yang, Weihang Zhang, Huazhu Fu, Huiqi Li

机构 * Beijing Institute of Technology, Beijing, China(北京理工大学) National University of Singapore, Singapore(新加坡国立大学) Tsinghua University, Beijing, China(清华大学) The Hong Kong Polytechnic University, Hong Kong(香港理工大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.AI

AI总结 本文提出EyExIn框架,通过深度专家注入机制提升视网膜VLMs的领域知识嵌入能力,解决感知与推理间隙问题,实现高精度眼科视觉问答。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24459 2026-07-29 cs.AI 版本更新 77%

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

从执行到能力:通过过程知识合成进行科学经验巩固

Liwei Dong, Jiahao Zhao, Nan Xu

机构 * XScience Lab(X科学实验室) Wenge AI(文阁人工智能)

专题命中 领域大模型 :large language model(abstract);language model(abstract);SFT(abstract);分类 cs.AI

AI总结 研究科学计算经验巩固问题,引入SciConsolidate方法,通过对比成败归纳跨任务过程,经开发验证门选择,用失败告知等扩展数据,克服抽象执行差距,在SciCode上实验,为科学计算建立经验到能力途径及自我改进起点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11882 2026-07-29 cs.CY cs.AI cs.HC cs.SE 77%

An Experience Report on a Pedagogically Controlled, Curriculum-Constrained AI Tutor for SE Education

关于在教学控制和课程约束下的AI导师在软件工程教育中的经验报告

Lucia Happe, Dominik Fuchß, Luca Hüttner, Kai Marquardt, Anne Koziolek

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 本文介绍了一种基于AI的导师系统,旨在通过结构化提示和语义标记知识库,为中学生提供个性化、课程约束的编程学习支持,并通过试点评估验证其在降低认知负荷和补充课堂教学中的潜力。

Comments 11 pages, 4 figures, accepted for publication at ICSE 2026 SEET Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02458 2026-07-28 cs.CY cs.AI cs.ET 版本更新 77%

Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments

当模拟看起来正确但因果效应出错:大型语言模型作为行为模拟器

Zonghan Li, Feng Ji

机构 * University of Toronto(多伦多大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究评估了大型语言模型在气候心理学干预中的行为模拟能力,发现描述性拟合不等于因果准确性,不同干预逻辑和结果类型导致误差差异,提示依赖描述性拟合可能误导干预效果结论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12650 2026-07-28 cs.CV cs.AI 版本更新 77%

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

AutoMat:通过智能工具使用实现从显微镜图像自动重建晶体结构

Yaotian Yang, Yiwen Tang, Yizhe Chen, Xiao Chen, Jiangjie Qiu, Hao Xiong, Haoyu Yin, Zhiyao Luo, Yifei Zhang, Sijia Tao, Wentao Li, Qinghua Zhang, Yuqiang Li, Wanli Ouyang, Bin Zhao, Xiaonan Wang, Fei Wei

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究旨在从有噪声的STEM投影重建晶体结构,提出AutoMat智能控制器,通过闭环验证和多模块组合实现。引入基准数据集评估,结果显示其性能优于现有方法,建立了从微观到原子尺度建模的途径。

Comments The code and dataset are publicly available at https://github.com/yyt-2378/AutoMat and https://huggingface.co/datasets/yaotianvector/STEM2Mat

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21887 2026-07-27 cs.HC cs.CL 新提交 77%

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

使用水平适配的具身对话代理降低外语焦虑

Krishan Rajaratnam, Wenbin Gan, Yuan Sun

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究针对外语焦虑影响二语习得问题,提出基于欧洲共同语言参考标准的多智能体具身对话系统,通过“生成-评估-再生”循环适配用户水平。小样本试点研究表明该系统生成的对话句子更适配学习者,虽未显著降低焦虑,但提供了相关见解。

Comments 8 pages, 6 figures, published in the proceedings of EDULEARN26

Journal ref K. Rajaratnam, W. Gan, Y. Sun (2026) TOWARDS REDUCING FOREIGN LANGUAGE ANXIETY USING LEVEL-APPROPRIATE EMBODIED CONVERSATIONAL AGENTS, EDULEARN26 Proceedings, Article 1459

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27347 2026-07-27 cs.CL 版本更新 77%

Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline

用多语言联合实体关系抽取管道绘制欧洲政治精英网络

Kirill Solovev, Jana Lasser

机构 * IDea_Lab, University of Graz(格拉茨大学IDea_Lab)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出模块化、全开放权重的多语言联合实体关系抽取管道,从大规模非结构化新闻语料中构建带符号的时间知识图谱,通过跨度命名实体识别、三阶段链接级联和约束混合专家模型实现高文本正确性,并在奥地利和波兰案例中验证其有效性。

Comments Updated NET-ROL citation per their request

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17902 2026-07-21 cs.DL cs.CL 新提交 77%

Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field

用于生物医学领域研究主题本体生成的资源高效语言模型基准测试

Tanay Aggarwal, Angelo Salatino, Francesco Osborne, Enrico Motta

机构 * Knowledge Media Institute, The Open University, Milton Keynes, UK(开放大学知识媒体研究所) Department of Business and Law, University of Milano-Bicocca, Milan, IT(米兰-比科卡大学商业与法律系)

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 本文评估五个小型开源LLMs识别生物医学概念语义关系的性能,引入MeSH-Rel-4K数据集并分析三种策略,发现针对性微调使平均F1分数显著提高,突破推理瓶颈,为构建生物医学本体提供准确自动化方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16967 2026-07-16 cs.AI cs.IR 77%

Empowering Medical Equipment Sustainability in Low-Resource Settings: An AI-Powered Diagnostic and Support Platform for Biomedical Technicians

赋能低资源环境下的医疗设备可持续性:一个由AI驱动的诊断与支持平台,用于生物医学技术人员

Bernes Lorier Atabonfack, Ahmed Tahiru Issah, Mohammed Hardi Abdul Baaki, Clemence Ingabire, Tolulope Olusuyi, Maruf Adewole, Udunna C. Anazodo, Timothy X Brown

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校) Medical Artificial Intelligence Laboratory(医学人工智能实验室) University of Pennsylvania(宾夕法尼亚大学) McGill University(麦吉尔大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究提出一个AI驱动的平台,帮助生物医学技术人员实时诊断和修复医疗设备,通过集成大语言模型和用户友好的界面,提高低资源环境下的医疗设备可持续性。

Comments Accepted at the MIRASOL Workshop at MICCAI 2025. To appear in Lecture Notes in Computer Science (LNCS)

Journal ref In: Anazodo U. et al. (eds) Medical Image Computing in Resource Constrained Settings. MIRASOL 2025. LNCS, Springer, Cham, pp. 217-230

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11074 2026-07-14 cs.CL 新提交 77%

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

ResearchQA:科学论文中基于引用的问答基准测试

Saba Imran, Debanjum Singh Solanky

机构 * Khoj Inc.(Khoj公司)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 该研究引入ResearchQA基准测试,含多领域多类型问答对,用于科学论文基于引用的问答评估。通过特定方法评估八个模型,发现基于引用指标区分度更高,开放权重模型接近最佳封闭模型准确率且延迟更低,还发布了相关资源。

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03174 2026-07-14 cs.AI 版本更新 77%

InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation

InqEduAgent:具有高斯过程增强的自适应人工智能学习伙伴

Wen-Xi Yang, Tian-Fang Zhao, Guan Liu

机构 * Guangdong Institute of Smart Education(广东智能教育研究院) Jinan University(暨南大学) School of Journalism and Communication(新闻与传播学院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究针对探究式教育中学习伙伴分配问题,提出InqEduAgent框架,集成高斯过程增强匹配机制,能基于先验知识模式选自适应伙伴,实验证明其性能优越,推进了人机协作学习及相关领域发展。

Comments Accepted by the 2026 8th Asia Conference on Machine Learning and Computing (ACMLC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07915 2026-07-10 cs.CY cs.CL 新提交 77%

Validating LLMs in social science: Epistemic threats and emerging norms

在社会科学中验证大语言模型:认知威胁与新兴规范

Meera Desai, Dallas Card, Abigail Z. Jacobs

机构 * School of Information, University of Michigan(信息学院,密歇根大学) Center for the Study of Complex Systems, University of Michigan(复杂系统研究中心,密歇根大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究探讨大语言模型对社会科学方法论的重塑,通过分析八本旗舰期刊论文语料库,发现其生成测量在实证分析中核心作用凸显,但验证实践存问题,进而概述补充策略以完善大语言模型在社会科学应用的规范标准。

Comments 28 pages, 2 figures. Main text: 11 pages, Appendix: 11 pages, References: 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19458 2026-07-08 cs.AI cond-mat.mtrl-sci 版本更新 77%

VASP Agent: An Agentic Framework for Autonomous First-principles Calculations

VASP智能体:一种用于自主第一性原理计算的智能框架

Zeyu Xia, Jinzhe Ma, Congjie Zheng, Zhongyao Wang, Shufei Zhang, Yuqiang Li, Hang Su, P. Hu, Changshui Zhang, Xingao Gong, Wanli Ouyang, Lei Bai, Dongzhan Zhou, Mao Su

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Computer Science and Technology(计算机科学与技术系) School of Physical Science and Technology(物理科学与技术学院) Department of Automation(自动化系) Beijing National Research Center for Information Science and Technology (BNRist)(北京信息科学与技术国家研究中心) School of Chemistry and Chemical Engineering(化学与化工学院) Key Laboratory of Computational Physical Sciences (Ministry of Education)(计算物理科学重点实验室) The Chinese University of Hong Kong(香港中文大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究针对第一性原理材料计算对自主性的高要求,提出VASP智能体系统,结合多种技能和工具执行多步VASP计算,经多任务评估,其计算结果更优,还能诊断并恢复计算错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04907 2026-07-07 cs.AI 新提交 77%

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation

Medi-Gemma:一种集成确定性电子病历分析与检索增强生成的混合临床决策支持系统

Mohammed Saim Ahmed Quadri, Yunzhe Xue, Justin W. Ady, Usman Roshan

机构 * New Jersey Institute of Technology(新泽西理工学院) Robert Wood Johnson Hospital(罗伯特·伍德·约翰逊医院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 介绍用于伤口病理分类和工作流程自动化的临床决策支持系统Medi-Gemma,其采用解耦框架,经多阶段管道处理数据请求,关键贡献是地面真值注入模块,验证表明该架构有诸多优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20698 2026-07-07 cs.CV cs.CL 版本更新 77%

Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs

多模态大语言模型在胃肠诊断中的临床认知对齐

Huan Zheng, Yucheng Zhou, Tianyi Yan, Dubing Chen, Hongbo Lu, Wenlong Liao, Tao He, Pai Peng, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学协同创新研究院物联网国家重点实验室) Shanghai Jiao Tong University(上海交通大学) COWARobot Co. Ltd.(COWARobot有限公司)

专题命中 领域大模型 :large language model(abstract);language model(abstract);SFT(abstract);分类 cs.CL

AI总结 本文提出CogAlign框架,通过构建层次化临床认知数据集和监督微调提升模型临床分析能力,并采用反事实强化学习消除视觉偏差,实现胃肠诊断的因果关联与高准确率。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20666 2026-07-07 cs.HC cs.CL 版本更新 77%

TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews

TAMA:一种使用多智能体大语言模型进行临床访谈的人机协作主题分析框架

Huimin Xu, Seungjun Yi, Terence Lim, Jiawei Xu, Andrew Well, Carlos Mery, Aidong Zhang, Yuji Zhang, Heng Ji, Keshav Pingali, Yan Leng, Ying Ding

机构 * School of Information, University of Texas at Austin(信息学院,德克萨斯大学奥斯汀分校) Department of Biomedical Engineering, University of Texas at Austin(生物医学工程系,德克萨斯大学奥斯汀分校) College of Natural Sciences, University of Texas at Austin(自然科学院,德克萨斯大学奥斯汀分校) Graphen, Inc.(Graphen公司) Department of Cardiac Surgery, Division of Pediatric Cardiac Surgery, Vanderbilt University School of Medicine(心脏外科系,范德比尔特大学医学中心) Pediatric Heart Institute, Monroe Carell Jr. Children’s Hospital at Vanderbilt(儿童心脏研究所,范德比尔特儿童医院) Department of Computer Science, University of Virginia(计算机科学系,弗吉尼亚大学) Department of Computer Science, University of Illinois at Urbana-Champaign(计算机科学系,伊利诺伊大学厄巴纳-香槟分校) Department of Computer Science, University of Texas at Austin(计算机科学系,德克萨斯大学奥斯汀分校) McCombs School of Business, University of Texas at Austin(麦克阿瑟商学院,德克萨斯大学奥斯汀分校) Dell Medical School, University of Texas at Austin(德克萨斯大学奥斯汀分校德莱尔医学院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出人机协作主题分析框架TAMA,利用多智能体系统结构化对话及心脏专家专业知识,用于临床访谈分析,在罕见病访谈转录本分析中性能优于单智能体方法。

Comments Manuscript accepted to ACM Transactions on Computing for Healthcare

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01960 2026-07-03 cs.CL 新提交 77%

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

NAVER LABS Europe 对指令遵循 2026 短轨道的提交

Marcely Zanon Boito, Hemant Yadav, Jean-Luc Meunier, Ioan Calapodescu

机构 * NAVER LABS Europe IIIT Delhi(印度德里国际信息技术学院)

专题命中 领域大模型 :LLM(abstract,abstract_cn);prompting(abstract);分类 cs.CL

AI总结 提出SpeechMapper语音投影器和合成SQA数据集fakACL,在更紧凑模型上实现ASR、ST和SQA联合处理,获得IWSLT 2026短轨道并列第一。

Comments IWSLT 2026 system paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01440 2026-07-03 cs.CL 新提交 77%

FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning

FaithMed: 训练大语言模型进行忠实于证据的医学推理

Zhiyun Zhang, Liwen Sun, Xiang Qian, Chenyan Xiong

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University School of Medicine(斯坦福大学医学院) Xlue

专题命中 领域大模型 :LLM(summary_cn,abstract_cn);分类 cs.CL

AI总结 提出FaithMed框架,通过过程级奖励和优势分组强化学习,监督LLM在医学推理中忠实评估和应用检索证据,在7个基准上平均提升9%。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26957 2026-07-03 cs.CY cs.AI 77%

Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings

模拟有效性:在MLLM生成的科学图表反馈中的模态解耦

Arne Bewersdorff, Nejla Yuruk, Xiaoming Zhai

机构 * University of Georgia, AI4STEM Education Center(佐治亚大学AI4STEM教育中心) Gazi University, Department of Mathematics and Science Education(加齐大学数学与科学教育系)

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 研究探讨了多模态大语言模型生成的科学图表反馈中模态解耦问题,发现反馈常存在 grounding 失败,表明需超越常规提示策略的 grounding 机制。

Comments Accepted as AIED Short Paper 2026, Seoul, South Korea. Submission #1147. This is the long paper version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28666 2026-06-30 cs.CR cs.AI 77%

Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare

为何信任你的智能体?医疗保健中TRiSM引导的智能体工作流的实证安全增益

Liam Kearns

机构 * AuraQ

专题命中 领域大模型 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文应用AI信任、风险与安全管理(TRiSM)框架,将不安全的医疗报告生成智能体工作流转变为安全敏感的工作流,在五种大语言模型上评估,将平均攻击成功率从31%-42%降至10%-25%,同时报告准确率提升14个百分点。

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21530 2026-06-30 cs.LG 77%

Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

通过基于查询的模型协作进行专家指导的临床文本增强

Dongkyu Cho, Miao Zhang, Rumi Chunara

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出基于查询的模型协作框架,利用专家知识指导文本增强,提升医疗信息保留并减少幻觉,实验显示在临床预测任务中表现优于传统方法。

Comments 18 pages, 6 figures, Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏