arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-07-28 至 2025-07-28 共收录 22 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 22 篇

2507.18143 2025-07-28 cs.CL cs.AI 90%

HIVMedQA: Benchmarking large language models for HIV medical decision support

Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Roger Kouyos, Niko Beerenwinkel, Diane Duroux

机构 * Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland(1 生物系统科学与工程系,苏黎世联邦理工学院,巴塞尔,瑞士) School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland(2 生命科学学院,洛桑联邦理工学院,洛桑,瑞士) Swiss Institute of Bioinformatics, Lausanne, Switzerland(3 瑞士生物信息学研究所,洛桑,瑞士) Biomedical Data Science Center, Lausanne University Hospital and University of Lausanne, Lausanne, Switzerland(4 生物医学数据科学中心,洛桑大学医院和洛桑大学,洛桑,瑞士) Institute of Medical Virology, University of Zurich, Zurich, Switzerland(5 医学病毒学研究所,苏黎世大学,苏黎世,瑞士) Department of Infectious Diseases and Hospital Epidemiology, University Hospital Zurich, Zurich, Switzerland(6 感染病与医院流行病学系,苏黎世大学医院,苏黎世,瑞士) ETH AI Center, ETH Zurich, Zurich, Switzerland(7 ETH人工智能中心,苏黎世联邦理工学院,苏黎世,瑞士) Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland(8 量化生物医学系,苏黎世大学,苏黎世,瑞士)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18119 2025-07-28 cs.CL cs.AI cs.SD eess.AS 88%

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Hongjie Chen, Zehan Li, Yaodong Song, Wenming Deng, Yitong Yao, Yuxin Zhang, Hang Lv, Xuechao Zhu, Jian Kang, Jie Lian, Jie Li, Chao Wang, Shuangyong Song, Yongxiang Li, Zhongjiang He, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究所(TeleAI),中国电信,中国)

专题命中 评测与基准 :language model(title,abstract);SLM(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19185 2025-07-28 cs.CR cs.AI 88%

PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models

Tarek Gasmi, Ramzi Guesmi, Mootez Aloui, Jihene Bennaceur

机构 * University of Manouba(曼努巴大学) University of Jendouba(杰努布大学) LETI Laboratory, University of Sfax(萨克斯大学LETI实验室) DataDoIt(DataDoIt公司) South Mediterranean University(南地中海大学)

专题命中 评测与基准 :language model(title,abstract);large language model(title);LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08088 2025-07-28 cs.CR cs.IR 88%

KGV: Integrating Large Language Models with Knowledge Graphs for Cyber Threat Intelligence Credibility Assessment

Zongzong Wu, Fengxiao Tang, Ming Zhao, Yufeng Li

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13666 2025-07-28 cs.CL cs.AI cs.CY 86%

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation

Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopal, Katarina Marcincinova, Matus Mesarcik

机构 * Kempelen Institute of Intelligent Technologies(凯普勒智能技术研究所) University of Copenhagen(哥本哈根大学) Comenius University in Bratislava(布拉迪斯拉瓦科เมนius大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 main

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19399 2025-07-28 cs.CR cs.AI 85%

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security

Gabriel Chua

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18957 2025-07-28 cs.SE 85%

SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents

Jianming Chang, Jieke Shi, Yunbo Lyu, Xin Zhou, Lulu Wang, Zhou Yang, Bixin Li, David Lo

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19004 2025-07-28 cs.CV cs.AI 79%

MedIQA: A Scalable Foundation Model for Prompt-Driven Medical Image Quality Assessment

Siyi Xun, Yue Sun, Jingkun Chen, Zitong Yu, Tong Tong, Xiaohong Liu, Mingxiang Wu, Tao Tan

机构 * Faculty of Applied Sciences, Macao Polytechnic University, Macao, China(澳门理工学院) Department of Engineer Science, University of Oxford, Oxford, UK(牛津大学工程科学系) Great Bay University, Dongguan, China(东莞大湾大学) College of Physics and Information Engineering, Fuzhou University, Fuzhou, China(福州大学物理与信息工程学院) Shanghai Jiao Tong University, Shanghai, China(上海交通大学) Department of Radiology, Shenzhen People’s Hospital, Shenzhen, China(深圳人民医院放射科)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

Comments We note that the version after peer review of this paper has been provisionally accepted by The 28th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19370 2025-07-28 cs.CV 78%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 评测与基准 :LLM(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19242 2025-07-28 cs.RO 78%

Foundation Model-Driven Grasping of Unknown Objects via Center of Gravity Estimation

Kang Xiangli, Yage He, Xianwu Gong, Zehan Liu, Yuru Bai

机构 * School of Electronics and Control Engineering, Chang’an University(电子与控制工程学院,长安大学)

专题命中 评测与基准 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15850 2025-07-28 cs.CL 77%

3LM: Bridging Arabic, STEM, and Code through Benchmarking

Basma El Amel Boussaha, Leen AlQadi, Mugariya Farooq, Shaikha Alsuwaidi, Giulia Campesan, Ahmed Alzubaidi, Mohammed Alyafeai, Hakim Hacid

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00151 2025-07-28 cs.CL cs.AI 73%

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, Abdelrahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-Chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar Al-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Rouabhia Anfel, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed

机构 * The University of British Columbia(不列颠哥伦比亚大学) MBZUAI(穆扎伊布人工智能研究所) Invertible AI(可逆人工智能) Birzeit University(比尔泽特大学) Prince Sultan University(沙特王储大学) UM6P Cairo University(开罗大学) JUST(朱拉大学) Ain Shams University(艾因·夏姆斯大学) Damascus University(大马士革大学) University of Khartoum(喀土穆大学) Menoufiya University(蒙努菲亚大学) University of Nouakchott(努瓦克肖特大学) National Polytechnic School of Algiers(阿尔及利亚国家多科技术学校) Full Sail University(全帆大学) Alfaisal University(阿尔法西大学) Hamad Bin Khalifa University(哈利法大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments More information about our dataset is available at our project page: https://github.com/UBC-NLP/palm

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06740 2025-07-28 cs.AI 70%

AI PsyRoom: Artificial Intelligence Platform for Segmented Yearning and Reactive Outcome Optimization Method

Yigui Feng, Qinglin Wang, Ke Liu, Xinhai Chen, Bo Yang, Jie Liu

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments I found that some of the experiments were wrong with some data, especially those involving the protocol evaluation area

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06303 2025-07-28 cs.CL cs.CV 70%

Long-Form Answers to Visual Questions from Blind and Low Vision People

Mina Huh, Fangyuan Xu, Yi-Hao Peng, Chongyan Chen, Hansika Murugu, Danna Gurari, Eunsol Choi, Amy Pavel

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Carnegie Mellon University(卡内基梅隆大学) Hong Kong University of Science and Technology(香港科学与技术大学) University of Colorado Boulder(科罗拉多大学博尔德分校)

专题命中 评测与基准 :language model(abstract);prompting(abstract);分类 cs.CL

Comments COLM 2024 Oral Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07510 2025-07-28 cs.AI cs.CR 70%

Secret Collusion among AI Agents: Multi-Agent Deception via Steganography

Sumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina, Philip H. S. Torr, Lewis Hammond, Christian Schroeder de Witt

机构 * UC Berkeley(加州大学伯克利分校) University of Oxford(牛津大学) Armasuisse Science+Technology(阿玛苏斯科学与技术) Google DeepMind(谷歌DeepMind)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17543 2025-07-28 cs.HC 67%

Anticipate, Simulate, Reason (ASR): A Comprehensive Generative AI Framework for Combating Messaging Scams

Xue Wen Tan, Kenneth See, Stanley Kok

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Comments arXiv admin note: text overlap with arXiv:2412.13528

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19315 2025-07-28 cs.CL cs.AI eess.AS 62%

JCAPT: A Joint Modeling Approach for CAPT

Tzu-Hsuan Yang, Yue-Yang He, Berlin Chen

机构 * National Taiwan Normal University(台湾师范大学)

专题命中 评测与基准 :prompting(abstract);分类 cs.CL、cs.AI

Comments Accepted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19470 2025-07-28 cs.CL cs.HC 57%

Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models

Son Quoc Tran, Tushaar Gangavarapu, Nicholas Chernogor, Jonathan P. Chang, Cristian Danescu-Niculescu-Mizil

机构 * Cornell University(康奈尔大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Harvey Mudd College(哈维·穆德学院)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments Code and data available as part of ConvoKit: https://convokit.cornell.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19398 2025-07-28 cs.CV cs.AI 57%

CXR-CML: Improved zero-shot classification of long-tailed multi-label diseases in Chest X-Rays

Rajesh Madhipati, Sheethal Bhat, Lukas Buess, Andreas Maier

机构 * Friedrich-Alexander University Erlangen-Nuremberg(弗里德里希-亚历山大大学埃尔朗根-纽伦堡)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18915 2025-07-28 cs.CL cs.CV 57%

Mining Contextualized Visual Associations from Images for Creativity Understanding

Ananya Sahu, Amith Ananthram, Kathleen McKeown

机构 * Columbia University(哥伦比亚大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18755 2025-07-28 cs.SE cs.AI cs.PL 57%

Agentic Program Repair from Test Failures at Scale: A Neuro-symbolic approach with static analysis and test execution feedback

Chandra Maddila, Adam Tait, Claire Chang, Daniel Cheng, Nauman Ahmad, Vijayaraghavan Murali, Marshall Roch, Arnaud Avondet, Aaron Meltzer, Victor Montalvao, Michael Hopko, Chris Waterson, Parth Thakkar, Renuka Fernandez, Kristian Kristensen, Sivan Barzily, Sherry Chen, Rui Abreu, Nachiappan Nagappan, Payam Shodjai, Killian Murphy, James Everingham, Aparna Ramani, Peter C. Rigby

机构 * Meta

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 50%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 评测与基准 :language model(abstract)

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏