arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-18 至 2025-11-18 共收录 428 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 94 篇

2511.13341 2025-11-18 cs.SE cs.AI 83%

An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains

Zihe Yan, Kai Luo, Haoyu Yang, Yang Yu, Zhuosheng Zhang, Guancheng Li

机构 * Zihe Yan 1(燕子和1) Kai Luo 2(罗凯2) Haoyu Yang 3(杨浩远3) Yang Yu 3(于洋3) Zhuosheng Zhang 1(张卓生1) Guancheng Li 3(李冠成3)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

Comments 7 figures, 4 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13225 2025-11-18 cs.CL 83%

Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms

Tyler Loakman, Joseph James, Chenghua Lin

机构 * Department of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学系) Department of Computer Science, University of Manchester, UK(曼彻斯特大学计算机科学系)

专题命中 评测与基准 :language model(title,abstract);large language model(abstract);分类 cs.CL

Comments Accepted to IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09724 2025-11-18 cs.AI 83%

UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge

Yang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu, Yidong Wang, Guangsheng Bao, Jie Tang

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10027 2025-11-18 cs.AI 83%

ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response

Risha Surana, Qinyuan Ye, Swabha Swayamdipta

机构 * University of Southern California(南加州大学)

专题命中 评测与基准 :language model(title,abstract);LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21424 2025-11-18 cs.CL cs.AI cs.CV cs.LG 82%

Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings

Abderrazek Abid, Thanh-Cong Ho, Fakhri Karray

机构 * MBZUAI University of Waterloo(滑铁卢大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref In: Rojas I., Ortuño F., Rojas Ruiz F., Herrera L.J., Valenzuela O., Escobar J.J. (eds) Bioinformatics and Biomedical Engineering. IWBBIO 2026. Lecture Notes in Bioinformatics, vol 15561. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11594 2025-11-18 cs.CL cs.AI 81%

TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy

James McCammon

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23639 2025-11-18 cs.LG cs.AI q-bio.QM 81%

Integrating Genomics into Multimodal EHR Foundation Models

Jonathan Amar, Edward Liu, Alessandra Breschi, Liangliang Zhang, Pouya Kheradpour, Sylvia Li, Lisa Soleymani Lehmann, Alessandro Giulianelli, Matt Edwards, Yugang Jia, David Nola, Raghav Mani, Pankaj Vats, Jesse Tetreault, T. J. Chen, Cory Y. McLean

机构 * Verily Life Sciences(Verily生命科学公司) Nvidia(英伟达公司) Google(谷歌公司)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00730 2025-11-18 cs.HC 80%

Teaching LLMs to See and Guide: Context-Aware Real-Time Assistance in Augmented Reality

Mahya Qorbani, Kamran Paynabar, Mohsen Moghaddam

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments This work has been submitted to the IEEE Transactions on Systems, Man, and Cybernetics: Systems for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12829 2025-11-18 cs.LG 79%

An Evaluation of Representation Learning Methods in Particle Physics Foundation Models

Michael Chen, Raghav Kansal, Abhijith Gandrakota, Zichun Hao, Jennifer Ngadiuba, Maria Spiropulu

机构 * Division of Physics, Mathematics and Astronomy(物理、数学和天文系) California Institute of Technology(加州理工学院) Particle Physics Division(粒子物理部) Fermi National Accelerator Laboratory(费米国家加速器实验室) Bexorg, Inc.(Bexorg公司)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12821 2025-11-18 cs.CL 79%

BioMedJImpact: A Comprehensive Dataset and LLM Pipeline for AI Engagement and Scientific Impact Analysis of Biomedical Journals

Ruiyu Wang, Yuzhang Xie, Xiao Hu, Carl Yang, Jiaying Lu

机构 * Emory University(埃默里大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10071 2025-11-18 cs.AI 79%

Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM

Mohd Ariful Haque, Justin Williams, Sunzida Siddique, Md. Hujaifa Islam, Hasmot Ali, Kishor Datta Gupta, Roy George

机构 * Clark Atlanta University(克拉克亚特兰大大学) Daffodil International University(达福尔国际大学) Ahsanullah University of Science and Technology(阿沙努拉大学科学与技术学院)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

Journal ref 2025 IEEE International Conference on Service-Oriented System Engineering (SOSE), Tucson, AZ, USA, 21-24 July 2025, IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03294 2025-11-18 cs.CL cs.AI 79%

NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty

Leonidas Zotos, Ivo Pascal de Jong, Matias Valdenegro-Toro, Andreea Ioana Sburlea, Malvina Nissim, Hedderik van Rijn

机构 * Center for Language and Cognition, University of Groningen(语言与认知中心,格罗宁根大学) Bernoulli Institute, University of Groningen(伯努利研究所,格罗宁根大学) Department of Experimental Psychology, University of Groningen(实验心理学系,格罗宁根大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 2 figures, presented at ECAI 2025 at the 2nd International Workshop on AI in Society, Education and Educational Research (AISEER)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23183 2025-11-18 cs.CL cs.AI cs.HC 79%

Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement

Gabriele Sarti, Vilém Zouhar, Malvina Nissim, Arianna Bisazza

机构 * CLCG, University of Groningen(格罗宁根大学认知语言学与认知科学研究中心) ETH Zurich(苏黎世联邦理工学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL、cs.AI

Comments Under review. Code: https://github.com/gsarti/labl/tree/main/examples/unsup_wqe Metrics: https://huggingface.co/datasets/gsarti/unsup_wqe_metrics

Journal ref Proceedings of EMNLP (2025) 18320-18337

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13537 2025-11-18 cs.LG cs.AI 79%

Competence-Aware AI Agents with Metacognition for Unknown Situations and Environments (MUSE)

Rodolfo Valiente, Praveen K. Pilly

机构 * Intelligent Systems Center, HRL Laboratories(智能系统中心,HRL实验室)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments Replaced all references to "self-awareness" with the more accurate term "self-assessment"; Updated Figure 2; Added recent pertinent work from the cognitive computational neuroscience literature; Removed the non-apples-to-apples comparison with Dreamer-v3 for self-assessment; Added additional experiments to validate the role of accurate self-assessment in effective self-regulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11583 2025-11-18 cs.LG cs.AI cs.IR 79%

Parallel and Multi-Stage Knowledge Graph Retrieval for Behaviorally Aligned Financial Asset Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 3 figures, RAGE-KG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04486 2025-11-18 cs.SE 78%

EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits

Wayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal, Jenny Liang, Wei-Lin Chiang, Anastasios Nikolas Angelopoulos, Ion Stoica, Graham Neubig, Ameet Talwalkar, Chris Donahue

专题命中 评测与基准 :LLM(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12224 2025-11-18 cs.CR 78%

RulePilot: An LLM-Powered Agent for Security Rule Generation

Hongtai Wang, Ming Xu, Yanpei Guo, Weili Han, Hoon Wei Lim, Jin Song Dong

专题命中 评测与基准 :LLM(title,abstract)

Comments This paper has been accepted for publication at ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13626 2025-11-18 cs.AI 77%

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang, Kaoyan Lu, Shuai Lyu, Yifan Zhu, Ping Zong Junpeng Ding, Xinyu Liu, Qunlin Chen, Weiwei Qin, Yiran Shen, Jiayi Cen

专题命中 评测与基准 :large language model(abstract);language model(abstract);instruction tuning(abstract);分类 cs.AI

Comments 13 pages, 3 figures,The 40th Annual AAAI Conference on Artificial Intelligence(AAAI 2026),Paper has been accepted for a poster presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01668 2025-11-18 cs.AI 77%

Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics

Yueqing Xi, Yifan Bai, Huasen Luo, Weiliang Wen, Hui Liu, Haoliang Li

机构 * Department of Electronic Engineering, City University of Hong Kong(DongGuan)(香港城市大学(东莞)电子工程系) Department of Electronic Engineering, City University of Hong Kong(香港城市大学电子工程系)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03898 2025-11-18 cs.CL 77%

Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles

Nusrat Jahan Lia, Shubhashis Roy Dipta, Abdullah Khan Zehady, Naymul Islam, Madhusodan Chakraborty, Abdullah Al Wasif

机构 * University of Dhaka(达卡大学) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Cisco Systems(思科系统) BanglaLLM Maharishi International University(玛哈里希国际大学) Unityflow AI

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments Accepted to BLP at AACL-IJCNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13953 2025-11-18 cs.CL 77%

ReviewGraph: A Knowledge Graph Embedding Based Framework for Review Rating Prediction with Sentiment Features

A. J. W. de Vink, Natalia Amat-Lefort, Lifeng Han

机构 * LIACS, Leiden University(莱顿大学LIACS) LUMC, Leiden University(莱顿大学LUMC)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments Peer-reviewed and published version is in ICKG-2025 (The 16th IEEE International Conference on Knowledge Graphs, November 13-14, 2025, Limassol, Cyprus)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08181 2025-11-18 cs.IR cs.AI 77%

MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System

Seung Hwan Cho, Yujin Yang, Danik Baeck, Minjoo Kim, Young-Min Kim, Heejung Lee, Sangjin Park

机构 * Department of Industrial Data Engineering, Hanyang University, Republic of Korea(工业数据工程系,翰阳大学) School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea(跨学科工业研究学院,翰阳大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments 13 pages, 2 figures, Accepted at RDGENAI at CIKM 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12630 2025-11-18 cs.CL cs.AI 76%

Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing

Maoqi Liu, Quan Fang, Yang Yang, Can Zhao, Kaiquan Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beihang University(北航) State Key Laboratory of CNS/ATM(国家空管流量管理技术实验室) Aviation Data Communication Corporation(航空数据通信公司)

专题命中 评测与基准 :LLM(title);分类 cs.CL、cs.AI

Comments Accepted to Advanced Engineering Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13242 2025-11-18 cs.CV 75%

MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection

Junjie Wu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);instruction tuning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02206 2025-11-18 cs.CV 75%

Language-Enhanced Generative Modeling for Amyloid PET Synthesis from MRI and Blood Biomarkers

Zhengjie Zhang, Xiaoxie Mao, Qihao Guo, Shaoting Zhang, Qi Huang, Mu Zhou, Fang Xie, Mianxin Liu

机构 * organization= Shanghai Artificial Intelligence Laboratory , city= Shanghai , postcode= 200082 , country= China organization= School of Medicine, Xiamen University , city= Xiamen , state= Fujian , country= China organization= Department of Nuclear Medicine \& PET Center, Huashan Hospital, Fudan University , city= Shanghai , country= China organization= Department of Gerontology, Shanghai Jiao Tong University Affiliated Sixth People’s Hospital , city= Shanghai , country= China organization= Department of Computer Science, Rutgers University , city= New Brunswick , state= New Jersey , country= United States organization= Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences , city= Shenzhen , country= China

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract)

Comments 31 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15501 2025-11-18 cs.CL cs.AI cs.LG 75%

DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios

Yao Huang, Yitong Sun, Yichi Zhang, Ruochen Zhang, Yinpeng Dong, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) College of AI, Tsinghua University(清华大学人工智能学院) Shanghai Qi Zhi Institute(上海启智研究所) State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 28 pages, 17 figures, accepted by NeruIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15621 2025-11-18 cs.SE 75%

DSCodeBench: A Realistic Benchmark for Data Science Code Generation

Shuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun, Qihao Zhu, Jie M. Zhang

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19002 2025-11-18 cs.CV cs.AI cs.CL cs.LG 75%

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Ziqi Yin, Wentao Hu, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara

机构 * Waseda University(早稻田大学) CyberAgent, Inc.(CyberAgent公司) AI Shift, Inc.(AI Shift公司) Nara Institute of Science and Technology(奈良研究所)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12472 2025-11-18 cs.CL cs.AI 73%

Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing

Mengying Wang, Chenhui Ma, Ao Jiao, Tuo Liang, Pengjun Lu, Shrinidhi Hegde, Yu Yin, Evren Gurkan-Cavusoglu, Yinghui Wu

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments The 40th AAAI Conference on Artificial Intelligence (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18982 2025-11-18 cs.LG cs.AI 73%

PAX-TS: Model-agnostic multi-granular explanations for time series forecasting via localized perturbations

Tim Kreuzer, Jelena Zdravkovic, Panagiotis Papapetrou

机构 * Department of Computer(计算机系) Systems Sciences Stockholm University(系统科学学院 斯德哥尔摩大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏