arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12706 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 12706 篇

2311.11096 2023-11-21 eess.IV cs.CV 73%

On the Out of Distribution Robustness of Foundation Models in Medical Image Segmentation

Duy Minh Ho Nguyen, Tan Ngoc Pham, Nghiem Tuong Diep, Nghi Quoc Phan, Quang Pham, Vinh Tong, Binh T. Nguyen, Ngan Hoang Le, Nhat Ho, Pengtao Xie, Daniel Sonntag, Mathias Niepert

专题命中 领域大模型 :foundation model(title,comments)

Comments Advances in Neural Information Processing Systems (NeurIPS) 2023, Workshop on robustness of zero/few-shot learning in foundation models

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16313 2026-08-28 cs.AI cs.LG 版本更新 73%

Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences

学习预测、发现和推理高维事件序列

Hugo Math

机构 * Faculty of Applied Computer Science(应用计算机科学系) University of Augsburg(艾格堡大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出统一事件序列建模、因果发现和大语言模型的框架,用于高维事件流中的自动化故障诊断,解决传统方法在高维数据中的不足。

Comments PhD dissertation, 135 pages of main content, 201 pages in total

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25973 2026-08-27 cs.AI cs.LG 新提交 73%

SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

SciMIF:理解科学领域中的多模态指令遵循

Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出用于评估多模态大语言模型科学指令遵循能力的SciMIF基准,揭示模型在科学领域性能差异及规模提升未对应改善约束遵循等问题,填补相关评估空白。

Comments 21pages, 9 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25370 2026-08-27 cs.IR cs.AI cs.LG 新提交 73%

CRAMER: Control via Request-Aware Masking for Editing Recommenders

CRAMER:基于请求感知掩码的推荐模型编辑控制框架

Zhiyuan Julian Su, Naihe Feng, Zhen Luther Qin, Ga Wu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) Faculty of Computer Science, Dalhousie University(达尔豪斯大学计算机科学学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 CRAMER是一种通过请求感知掩码调制冻结骨干参数的推荐控制框架,在多个大规模数据集上优于现有基线,开销最小且可控性与跨域适配能力更强,为请求感知序列推荐建立了新范式。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25071 2026-08-27 cs.CL cs.AI cs.CY 新提交 73%

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

HealthBench-Psych:OpenAI的HealthBench的心理健康子集

Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous

机构 * Beth Israel Deaconess Medical Center(贝斯以色列女执事医疗中心)

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 该研究针对现有通用健康基准无法单独评估心理健康性能的问题,构建了HealthBench-Psych心理健康子集,评估20个前沿及开源模型的心理健康相关表现,发现模型排名一致性高并发布了相关资源。

Comments 21 pages (6-page main text plus appendices), 2 figures. Code: https://github.com/mindbench-ai/healthbench-psych Data: https://huggingface.co/datasets/mindbench-ai/healthbench-psych

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16899 2026-08-27 cs.LG cs.AI 版本更新 73%

Epistemic Memory: A Validity Layer for Self-Maintaining Intelligent Systems

认知记忆:自维持智能系统的有效性层

Pin-Han Ho, Limei Peng, Yiming Miao, Yan Jiao

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文针对AI记忆机制忽略知识有效性条件导致语义坐标漂移的问题,提出认知记忆作为有效性维护层,引入OBM架构,实验表明其可提升智能体在变化观测条件下的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00559 2026-08-27 cs.CL cs.AI 73%

Polish-English medical knowledge transfer: A new benchmark and results

Łukasz Grzybowski, Jakub Pokrywka, Michał Ciesiółka, Jeremi I. Kaczmarek, Marek Kubis

机构 * Adam Mickiewicz University(亚当·密茨凯维奇大学) Poznan University of Medical Sciences(波兹南医科大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 9042-9063, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01859 2026-08-27 cs.CL cs.AI cs.IR 73%

Optimizing Retrieval-Augmented Generation of Medical Content for Spaced Repetition Learning

Jeremi I. Kaczmarek, Jakub Pokrywka, Krzysztof Biedalak, Grzegorz Kurzyp, Łukasz Grzybowski

机构 * Poznań University of Medical Sciences(波兹南医科大学) Adam Mickiewicz University(亚当·密茨凯维奇大学) Wydawnictwo Naukowe PWN(PWN科学出版社) SuperMemo World(SuperMemo世界公司) Association for Research and Applications of Artificial Intelligence(人工智能研究与应用协会)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Journal ref Proceedings of the 17th International Conference on Computer Supported Education (CSEDU 2025), Vol. 2, pp. 174-186, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24304 2026-08-26 cs.CL cs.AI 新提交 73%

SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling

SENSESHIFT:基于编码器的掩码填充实现连续情感控制文本生成

Shahed Masoudian, Markus Frohmann, Emmanouil Karystinaios, Navid Rekabsaz, Markus Schedl

机构 * Johannes Kepler University Linz(约翰开普勒林茨大学) Thomson Reuters Labs(汤森路透实验室) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究提出基于编码器的框架SenseShift,通过双向注意力等技术实现细粒度句子级情感控制文本生成,在故事与评论生成任务上,相比大解码器基线模型,情感可控性更强且文本质量与鲁棒性更优。

Comments Paper Accepted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23666 2026-08-26 cs.AI cs.CL 新提交 73%

Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

门控激活引导:减少医学问答中的谄媚行为与幻觉

Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出门控激活引导方法,通过推理时干预结合幻觉与谄媚行为的针对性引导,在40亿参数模型上提升医学问答抗压力能力,表现接近千亿参数模型

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06268 2026-08-26 cs.CL cs.LG 版本更新 73%

MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs

MPIB:用于医疗提示注入攻击和大语言模型临床安全性的基准

Junhyeok Lee, Han Jang, Kyu Sung Choi

机构 * Seoul National University College of Medicine(首尔国立大学医学院) Seoul National University(首尔国立大学) Seoul National University College of Medicine, Seoul National University Hospital(首尔国立大学医学院、首尔国立大学医院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 MPIB是一个用于评估医疗提示注入攻击和大语言模型临床安全性的基准,通过测量临床危害事件率和攻击成功率来评估模型的安全性。

Comments 19 pages, 7 figures, 18 tables. Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23248 2026-08-25 cs.CL cs.AI 新提交 73%

Future Querying: Can LLMs Serve as Implicit Medical World Models?

未来查询:大语言模型能否作为隐式医学世界模型?

Siri Willems, James Butterworth, Lore Goetschalckx, Peter Vrancx, Philippe Modard, Elke Giets, Ludovic Denoyer

机构 * imec, AI-labs(imec人工智能实验室)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出未来查询范式,探究LLMs能否作为隐式医学世界模型,其框架可基于非结构化临床文档运行,经微调的小型开放权重模型性能接近专有系统,在相关数据集上验证了LLMs可捕捉临床动态。

Comments This paper is accepted at The 1st MICCAI Workshop on Medical World Models (MICCAI-2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21673 2026-08-25 cs.LG cs.AI 新提交 73%

SynEHR: Joint Modeling Inter-visit Temporal Evolution and Intra-visit Clinical Structure for Longitudinal EHR Synthesis

SynEHR:联合建模就诊间时间演化与就诊内临床结构的纵向电子健康记录合成模型

Ximiao Li, Lin Jiang, Rongchao Xu, Dahai Yu, Zhe He, Guang Wang

机构 * Florida State University(佛罗里达州立大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本研究提出SynEHR框架,通过两项创新模块优化纵向EHR合成,在真实数据集的多维度评估中优于现有最优模型,生成更具临床一致性与时间真实性的EHR数据。

Comments 12 pages. CIKM '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21409 2026-08-25 cs.CY cs.AI cs.CL 新提交 73%

Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?

法庭中的谄媚者:大型语言模型(LLMs)是否对司法权威与演变的法律标准脆弱?

Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究通过比较诊断框架对比LLMs在法律与医学领域的表现,发现法律LLMs对司法权威扰动更脆弱,过度信任权威虚假信息,模型规模会放大该问题。

Comments Please cite the definitive, peer-reviewed version of this article published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), edited by Maria Liakata et al., Association for Computational Linguistics, pp. 10865-10886, 2026. DOI: https://doi.org/10.18653/v1/2026.acl-long.497

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, Association for Computational Linguistics, 2026, pp. 10865-10886

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03536 2026-08-25 cs.CL cs.AI 版本更新 73%

GraphMed-LT: Patient-Specific Graph Memory with Latent Clinical Thought Refinement for Multi-Turn Medical Conversations

GraphMed-LT:面向多轮医疗对话的具有潜在临床思维优化的患者特定图记忆

Zhaohan Meng, Zaiqiao Meng, Siwei Liu, Hao Xu, Ke Yuan, Iadh Ounis

机构 * School of Computing Science, University of Glasgow, UK(格拉斯哥大学计算机科学学院) School of Natural and Computing Science, University of Aberdeen, UK(阿伯丁大学自然与计算科学学院)

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 该研究针对多轮医疗对话中临床证据分散的问题,提出GraphMed-LT方法,通过患者特定图记忆及潜在临床思维优化,在多轮医疗QA任务上较基线实现显著提升。

Comments Accepted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19201 2026-08-21 cs.CL cs.AI cs.IR q-bio.QM 新提交 73%

Automatic bioinformatic software named entity recognition from literature

从文献中自动识别生物信息学软件命名实体

Hao Xuan, Rithvij Pasupuleti, Ben Liu, Haishuo Sun, Jun Zhang, Zijun Yao, Cuncong Zhong

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出混合命名实体识别框架SNAIL,结合词汇与语义建模策略,在基准数据集及真实文献上识别生物信息学软件/数据库,性能优于现有方法,可用于大规模文献的资源元分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18937 2026-08-20 cs.CL cs.AI 新提交 73%

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

MedUAG:面向医学多模态模型的统一理解与生成框架

Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学) Tencent Jarvis Lab(腾讯Jarvis实验室)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对医学多模态领域缺乏统一理解生成框架的问题,本文构建了MedUAGCorpus数据集与MedUAGBench基准,开发了MedUAG模型,其在医学理解与生成任务中表现出色,为下一代医学多模态系统奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17947 2026-08-19 cs.AI cs.LG cs.NE 新提交 73%

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

通过程序搜索与持续抽象发现实现过程式内容元生成

Matthew Siper, Ahmed Khalifa, Julian Togelius

机构 * New York University(纽约大学) University of Malta(马耳他大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 该研究在推箱子等4款游戏中,提出持续抽象发现(CAD)方法,结合2×2实验,证实其可提升进化程序搜索的适应度,且学习到的库能被后续程序复用。

Comments Accepted for publication in IEEE Conference on Games 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17330 2026-08-19 cs.AI cs.CL cs.CY 新提交 73%

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

用于医疗咨询的大语言模型评估时机过晚:预 formulation 差距

Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang

机构 * Harvard T.H. Chan School of Public Health(哈佛陈曾熙公共卫生学院) Harvard Medical School(哈佛医学院) Australian Artificial Intelligence Institute, University of Technology Sydney(悉尼科技大学澳大利亚人工智能研究所) Raycaster Beth Israel Deaconess Medical Center(贝斯以色列女执事医疗中心)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究指出医疗咨询LLMs评估时机过晚的预 formulation 差距,通过实验发现指导条件可调整咨询流程与记录,建议直接以首次接触行为评估该差距。

Comments 17 pages, 3 tables. Code, cases, prompts, complete transcripts, and results: https://github.com/ningkko/preformulation-gap

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25459 2026-08-19 cs.CL cs.LG 版本更新 73%

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

SimulRAG:基于模拟器的检索增强生成(RAG)框架,用于将大型语言模型(LLMs)落地到长篇科学问答任务中

Haozhou Xu, Dongxia Wu, Matteo Chinazzi, Ruijia Niu, Rose Yu, Yi-An Ma

机构 * University of California San Diego(加州大学圣迭戈分校) Stanford University(斯坦福大学) Northeastern University(东北大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 针对LLMs在长篇科学问答中易幻觉的问题,本文提出SimulRAG框架,引入UE+SBA机制,发布相关基准,实验表明其信息量与事实性较基线分别提升30.4%、16.3%

Comments Haozhou Xu and Dongxia Wu are co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15424 2026-08-18 cs.MA cs.AI cs.LG 新提交 73%

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

ETHOS:面向临床多智能体系统的模块化伦理框架

Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。

Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12845 2026-08-14 cs.IR cs.AI cs.LG 新提交 73%

FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation

FSGR:缓解基于语义ID(SID)的生成式推荐中的令牌频率偏差

Yuchen Zheng, Sihan Xu, Jingwen Yang, Xiangrui Cai, Haiwei Zhang, Xiaojie Yuan

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 针对基于SID的生成式推荐中存在的令牌频率偏差问题,提出FSGR框架,通过优化SID构建与训练策略缓解偏差,在保持准确率的同时提升了约20%的基尼公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09538 2026-08-14 cs.CL cs.AI 版本更新 73%

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

TCS-BENCH:评估最先进生成式AI理论计算机科学研究能力的基准

Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensson, Silvio Lattanzi, Mislav Balunovic, Theophane Weber, Vahab Mirrokni

机构 * Google Research(谷歌研究院) CNRS(法国国家科学研究中心) IRIF(法国信息学研究所) Université Paris-Cité(巴黎城市大学) Princeton University(普林斯顿大学) Université Paris-Saclay(巴黎萨克雷大学) CEA(法国原子能和替代能源委员会) California Institute of Technology(加州理工学院) Google DeepMind(谷歌DeepMind)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究人员推出TCS-Bench基准,基于STOC、FOCS、SODA论文的定理证明任务评估LLMs,结合验证智能体,参考验证器在专家标注集准确率超90%,为评估生成式AI的TCS研究能力提供工具

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17228 2026-08-14 cs.AI cs.LG 版本更新 73%

MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence

MatchMiner-AI:癌症临床试验匹配的开源解决方案

Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz, Muhammad Umar Afzal, Arshad A. Mohammed, Ayub Umair, Huan He, Chueh Husan Hsu, Sarah Sammons, James Lindsay, Emily Mallaber, Harry R. Klein, Gufran Gungor, Matthew Galvin, Michael Deletto, Sabrina Y. Camp, Stephen C. Van Nostrand, James Provencher, Joyce Yu, Naeem Tahir, Jonathan Wischhusen, Olga Kozyreva, Taylor Ortiz, Hande Tuncer, Jad El Masri, Alys Malcolm, Tali Mazor, Ethan Cerami, Kenneth L. Kehl

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 MatchMiner-AI通过开源平台基于合成EHR数据实现隐私保护的临床试验匹配AI解决方案,显著提升患者与试验匹配的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11420 2026-08-13 cs.AI cs.CL 新提交 73%

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

社会思维链:一种基于医学鉴别诊断方法论的多智能体架构

Del Coburn, Scott Sanner, Dan Silver

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出基于医学鉴别诊断方法论的多智能体架构SCoT,通过多轮专家协作提升复杂病例诊断召回率,其优势无法通过单一推理实现。

Comments 14 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10505 2026-08-12 cs.AI cs.CL cs.CV 新提交 73%

RadFusion: Towards Threshold-Controllable Radiology Report Generation

RadFusion:迈向阈值可控的放射学报告生成

Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz

机构 * Microsoft(微软公司)

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 RadFusion框架为放射学报告生成赋予阈值可控性,融合多标签分类器与VQA报告生成器,提升诊断准确率,支持ROC分析,适配不同临床场景,更易通过监管审批。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10045 2026-08-12 cs.LG cs.AI 新提交 73%

Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons

从垃圾信息中提取有效信号:基于成对比较联合学习奖励与工作者可靠性

Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal, Avishek Ghosh

机构 * IIT Bombay(印度理工学院孟买分校)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出一种基于EM算法的方法,联合学习成对比较场景下的物品奖励与工作者可靠性,通过Polya-Gamma潜变量优化模型,在真实与合成数据集上验证了其对垃圾工作者的强鲁棒性。

Comments Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026), Amsterdam, Netherlands, August 17-21, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08849 2026-08-12 cs.CL cs.AI cs.DB cs.MA cs.SC 版本更新 73%

SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching

SatIR:可扩展的高召回率约束满足基于信息检索的临床试验匹配

Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam

机构 * Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Samueli Electrical and Computer Engineering, UCLA(UCLA Samueli电气与计算机工程系) Department of Computer Science and Informatics, Emory University(埃默里大学计算机科学与信息学系) Mayo Clinic(梅奥诊所)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SatIR通过将临床试验资格条件和摘要转化为形式约束,结合SMT、关系代数和大语言模型,提升了临床试验匹配的召回率和效率,优于基于相似度的基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08958 2026-08-11 cs.LG cs.AI q-bio.GN q-bio.QM 新提交 73%

Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods

思路搜索:用思路引导树搜索以探索多样化的科学方法

Xuefei Julie Wang, Hao Cui, Michael P. Brenner, Subhashini Venugopalan

机构 * California Institute of Technology(加州理工学院) Google Research(谷歌研究院) Harvard University(哈佛大学)

专题命中 领域大模型 :LLM(abstract_cn);prompting(abstract);分类 cs.AI、cs.LG

AI总结 该研究针对树搜索在科学方法探索中易陷局部最优的问题,提出Idea Search框架,通过整合动态思路库优化树搜索,在scRNA-seq批次整合任务上提升了性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05611 2026-08-07 cs.CL cs.LG 新提交 73%

FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities

FOCUS:解耦大语言模型中的专家角色以提升领域专家能力

Guanyu Wang, Zidi Zhang, Xu Chu

机构 * Peking University(北京大学) The University of Sydney(悉尼大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 FOCUS通过正交分解解耦LLMs的专家角色,结合专家门控模块与两阶段训练策略,在多领域基准上提升了角色控制的任务准确率,性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏