arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12597 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 12597 篇

2511.21104 2026-05-15 cs.LG cs.PL 81%

BRIDGE: Building Representations In Domain Guided Program Synthesis

BRIDGE: 在领域引导的程序合成中构建表示

Robert Joseph George, Carson Eisenach, Udaya Ghai, Dominique Perrault-Joncas, Anima Anandkumar, Dean Foster

机构 * California Institute of Technology(加州理工学院) Amazon(亚马逊)

专题命中 领域大模型 :LLM(abstract_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 BRIDGE通过多艺术ifacts程序合成框架提升Lean验证正确性,结合代码、规范和定理证明领域,提高生成效率和正确率。

Comments 41 pages, 10 figures, 3 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13706 2026-05-14 cs.CR cs.AI cs.CY cs.NI 81%

Identifying AI Web Scrapers Using Canary Tokens

通过信标令牌识别人工智能网络爬虫

Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger

机构 * Duke University(杜克大学) University of Pittsburgh(匹兹堡大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一种自动识别与大型语言模型相关的网络爬虫的方法,通过动态网站和信标令牌验证爬虫身份,实验表明能可靠识别多个未公开的爬虫。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12263 2026-05-13 cs.DL cs.AI 81%

Reconnecting Fragmented Citation Networks with Semantic Augmentation

用语义增强重新连接碎片化的引文网络

Vu Thi Huong, Annika Buchholz, Imene Khebouri, Thorsten Koch, Tim Kunt, Wolfgang Peters-Kottig, Tomasz Stompor, Janina Zittel

机构 * Digital Data and Information for Society, Science, and Culture, Zuse Institute Berlin(数字数据与信息社会、科学与文化,柏林祖布研究所) Institute of Mathematics, Vietnam Academy of Science and Technology(越南科学技术 academy 数学研究所) Software and Algorithms for Discrete Optimization, Technische Universität Berlin(离散优化软件与算法,柏林技术大学) Applied Optimization, Zuse Institute Berlin(应用优化,柏林祖布研究所) Kooperativer Bibliotheksverbund Berlin-Brandenburg (KOBV), Zuse Institute Berlin(柏林-勃兰登堡合作图书馆联合会 (KOBV),柏林祖布研究所)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一种结合引文拓扑和大语言模型的高效框架,通过添加语义边和文本相似性加权,减少引文网络碎片化,保留学科同质性,并提升聚类分析的结构可解释性。

Comments 11 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11280 2026-05-13 gr-qc astro-ph.HE cs.AI 81%

Discovery of Interpretable Surrogates via Agentic AI: Application to Gravitational Waves

通过代理AI发现可解释的替代模型:应用于引力波

Tousif Islam, Digvijay Wadekar, Tejaswi Venumadhav, Matias Zaldarriaga, Ajit Kumar Mehta, Javier Roulet, Barak Zackay

机构 * Kavli Institute for Theoretical Physics, University of California Santa Barbara(加州大学圣芭芭拉分校凯文利理论物理研究所) Center for Gravitational Physics, University of Texas at Austin(德克萨斯大学奥斯汀分校重力物理中心) Department of Physics, University of California at Santa Barbara(加州大学圣芭芭拉分校物理系) International Centre for Theoretical Sciences, Tata Institute of Fundamental Research(塔塔基础研究机构国际理论科学研究中心) School of Natural Sciences, Institute for Advanced Study(高级研究院自然科学学院) Chennai Mathematical Institute(钦奈数学研究所) Kavli Institute for Cosmological Physics, The University of Chicago(芝加哥大学凯文利宇宙物理研究所) Department of Particle Physics & Astrophysics, Weizmann Institute of Science(魏茨曼科学研究所粒子物理与天体物理系)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出GWAgent,一种基于大语言模型的流程,直接从模拟数据构建可解释的分析替代模型。通过物理指导的域假设提升模型精度,实现高精度、快速且可解释的引力波模拟替代。

Comments 25 pages, 9 figures, codes available at https://github.com/tousifislam/GWAgent

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11259 2026-05-13 cs.AI 81%

Template-as-Ontology: Configurable Synthetic Data Infrastructure for Cross-Domain Manufacturing AI Validation

模板作为本体:用于跨领域制造人工智能验证的可配置合成数据基础设施

Grama Chethan

机构 * Siemens Digital Industries Software(西门子数字工业软件)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一种模板作为本体的方法,通过单一Python配置模块生成制造仿真和AI分析工具的统一数据结构,验证了跨领域制造AI验证的可行性。

Comments 18 pages, 1 fugure

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11234 2026-05-13 cs.AI 81%

The Semantic Training Gap: Ontology-Grounded Tool Architectures for Industrial AI Agent Systems

语义训练鸿沟:面向工业AI代理系统的本体引导工具架构

Grama Chethan

机构 * Siemens Digital Industries Software(西门子数字工业软件)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出基于本体的工具架构,解决工业AI代理系统中语义训练鸿沟问题,通过类型关系配置增强语义约束,减少领域标识符的幻觉率,实现跨域配置能力。

Comments 29 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11221 2026-05-13 q-bio.QM cs.LG 81%

Beyond Manual Curation: Augmenting Targeted Protein Degradation Databases via Agentic Literature Extraction Workflows

超越手动整理:通过代理文献提取流程增强靶向蛋白降解数据库

Yaochen Rao, Farzaneh Jalalypour, N. M. Anoop Krishnan, Rocío Mercado

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Chalmers University of Technology(楚姆勒斯技术大学) University of Gothenburg(哥德堡大学) Yardi School of Artificial Intelligence(Yardi人工智能学院) Indian Institute of Technology Delhi(德里印度理工学院)

专题命中 领域大模型 :LLM(summary_cn,abstract);分类 cs.LG

AI总结 本文提出一种基于代理的LLM工作流,通过三角比较提升靶向蛋白降解数据库的提取效率,实现了数据库的自动化扩展与验证,提高了数据质量和应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10791 2026-05-12 cs.AI 81%

PathISE: Learning Informative Path Supervision for Knowledge Graph Question Answering

PathISE: 通过信息路径监督学习知识图谱问答

Shengxiang Gao, Chao Lei, Jey Han Lau, Jianzhong Qi

机构 * The University of Melbourne(墨尔本大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 PathISE通过学习高质量的中间监督信号提升知识图谱问答性能,利用轻量级Transformer估计关系路径信息,生成可复用的监督信号以增强现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10582 2026-05-12 cs.CR cs.AI 81%

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing

通过干扰-校正平滑实现保证的对抗防御

Zheng Lin, Zhenxing Niu, Haoxuan Ji, Haichang Gao

机构 * Xidian University(西安电子科技大学) Xi’an Jiaotong University(西安交通大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种基于平滑的新型防御方法,通过干扰-校正机制提升大语言模型对劫持攻击的防御能力,理论分析提供了防御成功的紧界和干扰强度要求,实验表明其在安全性和有用性上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10155 2026-05-12 cs.CL 81%

NyayaAI: An AI-Powered Legal Assistant Using Multi-Agent Architecture and Retrieval-Augmented Generation

NyayaAI:基于多智能体架构和检索增强生成的AI法律助理

Deepanshu, Divi Saxena, Deepali Rana, Ayesha Varshney, Sahinur Rahman Laskar

机构 * School of Computer Science UPES, Dehradun, India(计算机科学学院 UPES 德里胡迪恩印度)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出NyayaAI,通过多智能体架构和检索增强生成技术,提升法律信息获取效率,实现法律流程自动化与简化,实验显示其在法律分类、检索和响应准确性上达到70%-74%的精度。

Comments 3 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02678 2026-05-12 cs.LG cs.ET cs.HC stat.ME stat.ML 81%

Causal Discovery Should Embrace the Wisdom of the Crowd

因果发现应拥抱大众智慧

Ryan Feng Lin, Yuantao Wei, Huiling Liao, Xiaoning Qian, Shuai Huang

机构 * Department of Industrial and Systems Engineering, University of Washington(华盛顿大学工业与系统工程系) Department of Applied Mathematics, Illinois Institute of Technology(伊利诺伊理工学院应用数学系) Department of Electrical & Computer Engineering, Texas A&M University(德克萨斯阿灵顿大学电气与计算机工程系)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文探讨了因果学习中大众智慧范式的发展,提出通过分布式和基于群体的方法整合分散的因果知识,构建全局因果结构,并呼吁跨学科合作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12334 2026-05-12 cs.CL 81%

QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model

QM-ToT: 一种用于量化模型的医学树状思维推理框架

Zongxian Yang, Jiayu Qian, Kay Chen Tan, Hau-San Wong, Yulong Chen, Haoyu Zhang, Zhi-An Huang

机构 * City University of Hong Kong(Dongguan)(香港城市大学(东莞)) Hong Kong Polytechnic University(香港理工大学) City University of Hong Kong(香港城市大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出QM-ToT框架,通过树状思维方法分解复杂医疗问题,提升INT4量化模型在MedQAUSMLE数据集上的性能,使LLaMA2-70b模型准确率从34%提升至50%。

Comments Accepted by ICIC 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07572 2026-05-11 cs.AI stat.ML 81%

Open-Ended Task Discovery via Bayesian Optimization

通过贝叶斯优化进行开放任务发现

Masaki Adachi, Yuta Suzuki, Juliusz Ziomek

机构 * Lattice Lab Toyota Motor Corporation(电装株式会社拉特实验室) Machine Learning Research Group(机器学习研究组) University of Oxford(牛津大学)

专题命中 领域大模型 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 本文提出GSR框架,通过生成-选择-细化流程实现开放-ended的贝叶斯优化,应用于新产品开发、化学合成放大、算法分析和专利再利用,优于现有LLM优化器。

Comments 60 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06981 2026-05-11 cs.IR cs.CL 81%

Bridging Textual Profiles and Latent User Embeddings for Personalization

弥合文本特征与潜在用户嵌入之间的鸿沟以实现个性化

Zhaoxuan Tan, Xiang Zhai, Yan Zhu, Meng Jiang, Mohamed Hammad

机构 * University of Notre Dame(诺特大学) Google(谷歌)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出BLUE框架,通过结合语言基用户特征与嵌入基推荐目标,弥合可解释性文本特征与判别性潜在嵌入之间的差距,实验证明其在零样本序列推荐中优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06822 2026-05-11 cs.LG 81%

SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents

SHARP: 一种自进化可审计的规则策略用于金融交易代理

Xiwen Chen, Wenhui Zhu, Songzhu Zheng, Kashif Rasul, Yueyue Deng, Huayu Li

机构 * Morgan Stanley(摩根大通) Arizona State University(亚利桑那州立大学) Columbia University(哥伦比亚大学) University of Arizona(亚利桑那大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 SHARP通过结构化规则优化解决金融交易代理中信用分配问题,提升策略鲁棒性和透明度,使紧凑模型性能提升10-20个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06335 2026-05-08 cs.LG 81%

Eliciting associations between clinical variables from LLMs via comparison questions across populations

通过跨人群比较问题从LLMs中提取临床变量之间的关联

Fabian Kabus, Kian Kordtomeikel, Thomas Brox, Heinz Wiendl, Daiana Stolz, Harald Binder

机构 * Institute of Medical Biometry and Statistics (IMBI), Medical Center, University of Freiburg(弗赖堡大学医学生物统计学研究所(IMBI)、弗赖堡大学医学中心) Department of Computer Science, Faculty of Engineering, University of Freiburg(弗赖堡大学工程学院计算机科学系) Department of Pneumology, Medical Center, University of Freiburg(弗赖堡大学呼吸科医学中心) Department of Neurology and Neurophysiology, Medical Center, University of Freiburg(弗赖堡大学神经学与神经生理学医学中心)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出通过结构化比较问题从LLMs中提取临床变量关联,结合统计模型估计相关性,验证在COPD和MS领域中获得保守的因果预测候选链接。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27092 2026-05-01 cs.AI physics.optics 81%

End-to-end autonomous scientific discovery on a real optical platform

端到端自主科学发现于真实光学平台

Shuxing Yang, Fujia Chen, Rui Zhao, Junyao Wu, Yize Wang, Haiyao Luo, Ning Han, Qiaolu Chen, Yuze Hu, Wenhao Li, Mingzhu Li, Hongsheng Chen, Yihao Yang

机构 * State Key Laboratory of Extreme Photonics and Instrumentation, College of Information Science and Electronic Engineering, ZJU-Hangzhou Global Scientific and Technological Innovation Center, The Electromagnetics Academy at Zhejiang University, Zhejiang University(极端光子学与仪器国家重点实验室,信息科学与电子工程学院,浙大杭州国际科学技术创新中心,浙江大学电磁学学院,浙江大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Qiushi Discovery Engine,通过非线性研究阶段、Meta-Trace记忆和双层架构,在真实光学平台实现端到端自主科学发现,首次实验验证了非平凡物理机制。

Comments 25 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26959 2026-05-01 cs.CY cs.AI cs.MA 81%

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

CareGuardAI:面向临床安全与幻觉抑制的上下文感知多智能体守卫机制

Elham Nasarian, Abhilash Neog, Kwok-Leung Tsui, Niyousha HosseiniChimeh

机构 * Grado Department of Industrial & Systems Engineering, Virginia Tech, Blacksburg, VA 24061, USA(弗吉尼亚理工大学格拉多工业与系统工程系,弗吉尼亚理工大学,布莱克斯堡,VA 24061,美国) Department of Computer Science, Virginia Tech, Blacksburg, VA 24061, USA(弗吉尼亚理工大学计算机科学系,弗吉尼亚理工大学,布莱克斯堡,VA 24061,美国) Dept of Industrial, Manufacturing, and Systems Engineering at University of Texas at Arlington, Arlington, TX 76019, USA(德克萨斯理工大学阿灵顿分校工业、制造与系统工程系,阿灵顿,TX 76019,美国)

专题命中 领域大模型 :LLM(summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 CareGuardAI通过引入临床安全风险评估和幻觉风险评估,结合多阶段管道和迭代优化,提升患者面对LLM的可靠性与安全性,优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23949 2026-04-28 cs.AI 81%

Context-Aware Hospitalization Forecasting Evaluations for Decision Support using LLMs

基于上下文的医院化疗预测评估:利用LLMs进行决策支持

Rhea Makkuni, Ananya Joshi

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文评估了利用LLMs进行医院化疗预测的方法,通过三种方法在60个县的数据上验证了上下文增强的混合模型在非平稳医疗资源预测中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23801 2026-04-28 cs.CL cs.IR 81%

Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

领域微调与检索增强生成在医学多选问答中的比较:在4B参数规模下的受控比较

Avi-ad Avraam Buskila

机构 * Department of Information Science and Applied Artificial Intelligence(信息科学与应用人工智能系)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究比较了领域微调与检索增强生成在医学多选问答中的效果,发现领域微调在多数投票准确率上优于通用模型,而检索增强生成未显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23446 2026-04-28 cs.AI 81%

IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance

IndustryAssetEQA: 一种用于工业资产维护中具身体验问答的神经符号操作智能系统

Chathurangi Shyalika, Dhaval Patel, Amit Sheth

机构 * Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) University of South Carolina(南卡罗来纳大学) IBM Yorktown(IBM约克镇分公司) Indian AI Research Organization(印度人工智能研究组织)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出IndustryAssetEQA系统,结合事件 telemetry 表示与FMEA-KG知识图谱,提升工业资产问答的结构有效性、反事实准确性及解释蕴含性,减少专家评估的过度声称。

Comments 20 pages, 4 figures, 4 tables, Accepted for the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22002 2026-04-27 cs.CL 81%

When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation

当牛尿治愈便秘时:LLMs在检测文化特定健康谣言中的局限性

Anamta Khan, Ratna Kandala, Deepti, Sheza Munir, Joyojeet Pal

机构 * University of Michigan(密歇根大学) University of Kansas(堪萨斯大学) IIT Jodhpur(印度理工学院乔浦尔)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究通过分析印度YouTube上的牛尿相关内容,揭示LLMs在处理文化特定健康谣言时的局限性,指出文化嵌入的谣言无法通过提示工程解决。

Comments To appear in the proceedings of the 2nd Workshop on Misinformation Detection in the Era of LLMs (MisD), The 20th International AAAI Conference on Web and Social Media (ICWSM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20906 2026-04-24 cs.SE cs.AI 81%

Biomedical systems biology workflow orchestration and execution with PoSyMed

生物医学系统生物学工作流编排与执行与PoSyMed

Simon Süwer, Zoe Chervontseva, Kester Bagemihl, Jan Baumbach, Olga Tsoy, Andreas Maier

机构 * Institute for Computational Systems Biomedicine, University of Hamburg(计算系统生物医学研究所,汉堡大学) Faculty of Science, Computer Science, Vrije Universiteit Amsterdam(科学与计算机科学学院,阿姆斯特丹自由大学) Department for Mathematics and Computer Science, University of Southern Denmark(数学与计算机科学系,南丹麦大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 PoSyMed平台通过模块化架构和容器化流程,提升生物医学分析的可重复性与透明度,利用大语言模型辅助工具识别与参数设置。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20055 2026-04-23 cs.AI cs.HC 81%

From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI

从模糊到正式:利用AI扩展医院质量改进

Patrick Vossler, Jean Feng, Venkat Sivaraman, Robert Gallo, Hemal Kanzaria, Dana Freiser, Christopher Ross, Amy Ou, James Marks, Susan Ehrlich, Christopher Peabody, Lucas Zier

机构 * University of California, San Francisco(加州大学旧金山分校) Zuckerberg San Francisco General Hospital(扎克伯格旧金山总医院)

专题命中 领域大模型 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 本文提出通过AI框架优化医院质量改进过程,通过学习LLM提示和自然语言规范,提升质量因素发现的效率与可追溯性,实现人类与AI协同优化。

Comments 34 pages, 8 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25223 2026-04-23 cs.AI 81%

FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data

FELA:一种用于工业事件日志数据特征工程的多智能体进化系统

Kun Ouyang, Haoyu Wang, Dong Fang

机构 * Department of Electronic Engineering, Tsinghua University LIGHTSPEED STUDIOS, China

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出FELA,一种基于大语言模型的多智能体系统,通过协作生成、验证和实现新颖特征,提升工业事件日志数据的特征工程效率与性能。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03624 2026-04-23 cs.HC cs.CL 81%

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

LLAMADRS:在真实临床访谈中评估开源大语言模型——理性推理还是不理性推理?

Gaoussou Youssouf Kebe, Jeffrey M. Girard, Einat Liebenthal, Justin Baker, Fernando De la Torre, Louis-Philippe Morency

机构 * Carnegie Mellon University, School of Computer Science(卡内基梅隆大学计算机学院) University of Kansas, Department of Psychology(堪萨斯大学心理学系) McLean Hospital, Harvard Medical School(麦肯纳医院哈佛医学院)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文通过LLAMADRS基准测试,评估了25种开源大语言模型在结构化临床评估中的表现,发现理性推理模型在误差控制上并不总是更优,提示提示设计对模型性能有关键影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18373 2026-04-21 econ.GN cs.AI q-fin.EC q-fin.GN 81%

Dissecting AI Trading: Behavioral Finance and Market Bubbles

解析人工智能交易:行为金融与市场泡沫

Shumiao Ouyang, Pengfei Sui

机构 * Saïd Business School, University of Oxford(牛津大学said商学院) School of Management and Economics, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)管理学院)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究AI代理在实验资产市场中的预期形成与交易行为,发现AI代理表现出经典行为模式,并通过分析其推理文本展示针对性提示干预对市场泡沫的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14548 2026-04-21 cs.SD cs.LG eess.AS 81%

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

VoxSafeBench:不仅仅是所说的内容,还有谁、如何以及在哪里

Yuxiang Wang, Hongyu Liu, Yijiang Xu, Qinke Ni, Li Wang, Wan Lin, Kunyu Feng, Dekun Chen, Xu Tan, Lei Wang, Jie Shi, Zhizheng Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) Amphion Technology Co., Ltd.(Amphion科技有限公司)

专题命中 领域大模型 :SLM(summary_cn,abstract_cn);language model(abstract);分类 cs.LG

AI总结 VoxSafeBench首次联合评估SLM在安全、公平和隐私三个维度上的社会契合度,揭示语音识别中的普遍缺口,即模型在文本中能识别社会规范,但在语音中却无法正确应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06221 2026-04-21 cs.CL 81%

BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

BenchMarker:一种受教育启发的工具包,用于突出多项选择基准测试中的缺陷

Nishant Balepur, Bhavya Rajasekaran, Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber

机构 * University of Maryland(马里兰大学) New York University(纽约大学) Scale AI George Mason University(乔治·梅森大学)

专题命中 领域大模型 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 本文提出BenchMarker工具,利用LLM检测多项选择基准测试中的三项常见缺陷,通过人工标注和审计12个基准测试,发现自动和众包数据中存在大量缺陷,且修复方法可能引入新问题,影响NLP评估质量。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14609 2026-04-17 cs.AI physics.comp-ph 81%

El Agente Forjador: Task-Driven Agent Generation for Quantum Simulation

El Agente Forjador:面向量子模拟的任务驱动代理生成

Zijian Zhang, Aiwei Yin, Amaan Baweja, Jiaru Bai, Ignacio Gustin, Varinia Bernales, Alán Aspuru-Guzik

机构 * Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Department of Chemistry, University of Toronto(多伦多大学化学系) Department of Materials Science & Engineering, University of Toronto(多伦多大学材料科学与工程系) Department of Chemical Engineering & Applied Chemistry, University of Toronto(多伦多大学化学工程与应用化学系) Acceleration Consortium(加速联盟) Vector Institute for Artificial Intelligence(人工智能矢量研究所) Canadian Institute for Advanced Research (CIFAR)(加拿大高级研究 institute (CIFAR)) NVIDIA

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出El Agente Forjador框架,通过四阶段流程自主生成和复用计算工具,提升量子化学和动力学任务的准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏