arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-10-20 至 2025-10-20 共收录 160 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 37 篇

2510.15007 2025-10-20 cs.CL cs.AI 90%

Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective

Zhiqiang Kou, Junyang Chen, Xin-Qiang Cai, Ming-Kun Xie, Biao Liu, Changwei Wang, Lei Feng, Yuheng Jia, Gang Niu, Masashi Sugiyama, Xin Geng

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) RIKEN Center for Advanced Intelligence Project (AIP)(RIKEN先进人工智能项目中心) School of Computer Science and Technology, Qilu University of Technology(齐鲁大学计算机科学与技术学院) The University of Tokyo(东京大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03360 2025-10-20 cs.AI 89%

CogBench: A Large Language Model Benchmark for Multilingual Speech-Based Cognitive Impairment Assessment

Rui Feng, Zhiyao Luo, Wei Wang, Yuting Song, Yong Liu, Tingting Zhu, Jianqing Li, Xingyao Wang

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.AI

Comments 19 pages, 9 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15260 2025-10-20 cs.LG cs.AI cs.CL 89%

DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models

Yangyang Li

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) Massachusetts Institute of Technology(麻省理工学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint. Under review at ICLR 2026. 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15436 2025-10-20 cs.CL 88%

Controllable Abstraction in Summary Generation for Large Language Models via Prompt Engineering

Xiangchen Song, Yuchen Liu, Yaxuan Luan, Jinxu Guo, Xiaofan Guo

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17675 2025-10-20 cs.CL 88%

Evaluating Large Language Models with Psychometrics

Yuan Li, Yue Huang, Hongyi Wang, Ying Cheng, Xiangliang Zhang, James Zou, Lichao Sun

机构 * University of Cambridge(剑桥大学) University of Notre Dame(圣母大学) Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学) Lehigh University(莱斯利大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06153 2025-10-20 cs.SE cs.CL 86%

What's Wrong with Your Code Generated by Large Language Models? An Extensive Study

Shihan Dou, Haoxiang Jia, Shenxi Wu, Huiyuan Zheng, Muling Wu, Yunbo Tao, Ming Zhang, Mingxu Chai, Jessica Fan, Zhiheng Xi, Rui Zheng, Yueming Wu, Ming Wen, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang

机构 * Fudan University(复旦大学) Peking University(北京大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Shanghai Qiji Zhifeng Co., Ltd.(上海启智风科技有限公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments Accepted by SCIENCE CHINA Information Sciences (SCIS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15081 2025-10-20 cs.CL cs.SI 85%

A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling

Shiyu Ji, Farnoosh Hashemi, Joice Chen, Juanwen Pan, Weicheng Ma, Hefan Zhang, Sophia Pan, Ming Cheng, Shubham Mohole, Saeed Hassanpour, Soroush Vosoughi, Michael Macy

机构 * Cornell University(康奈尔大学) Georgia Institute of Technology(佐治亚理工学院) Dartmouth College(达特茅斯学院)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22467 2025-10-20 cs.MA cs.AI cs.LG 84%

Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems

Jiaxi Yang, Mengqi Zhang, Yiqiao Jin, Hao Chen, Qingsong Wen, Lu Lin, Yi He, Srijan Kumar, Weijie Xu, James Evans, Jindong Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) William & Mary(威廉与玛丽学院) Georgia Institute of Technology(佐治亚理工学院) AMD(AMD公司) Squirrel AI Amazon(亚马逊公司) University of Chicago(芝加哥大学)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03345 2025-10-20 cs.CR 83%

Elevating Cyber Threat Intelligence against Disinformation Campaigns with LLM-based Concept Extraction and the FakeCTI Dataset

Domenico Cotroneo, Roberto Natella, Vittorio Orbinato

专题命中 评测与基准 :LLM(title);large language model(abstract,comments);language model(abstract,comments)

Comments Accepted for publication in the Journal of Systems and Software (Special Issue on Reliable and Secure Large Language Models for Software Engineering)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15406 2025-10-20 cs.CL 83%

VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency

Hongcheng Liu, Yixuan Hou, Heyang Liu, Yuhao Wang, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15690 2025-10-20 cs.SE cs.CR 82%

MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing

Shiwen Ou, Yuwei Li, Lu Yu, Chengkun Wei, Tingke Wen, Qiangpu Chen, Yu Chen, Haizhi Tang, Zulie Pan

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract)

Comments Accepted for publication in IEEE Transactions on Software Engineering (TSE), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15706 2025-10-20 cs.IR cs.CL 81%

GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery

Italo Luis da Silva, Hanqi Yan, Lin Gui, Yulan He

机构 * Department of Informatics(信息学系)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments 9 pages, 6 figures, 3 tables, EMNLP 2025 Demo paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15685 2025-10-20 cs.CL 81%

Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection

Joshua Wolfe Brook, Ilia Markov

机构 * Computational Linguistics \& Text Mining Lab Vrije Universiteit Amsterdam

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments 8 pages, 9 figures, submitted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26490 2025-10-20 cs.CL cs.AI 81%

VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

Wei He, Yueqing Sun, Hongyan Hao, Xueyuan Hao, Zhikang Xia, Qi Gu, Chengcheng Han, Dengchang Zhao, Hui Su, Kefeng Zhang, Man Gao, Xi Su, Xiaodong Cai, Xunliang Cai, Yu Yang, Yunke Zhao

机构 * Meituan LongCat Team(美团LongCat团队)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

Comments The code, dataset, and leaderboard are available at https://vitabench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21826 2025-10-20 cs.CV cs.AI cs.LG 81%

Few-Shot Segmentation of Historical Maps via Linear Probing of Vision Foundation Models

Rafael Sterzinger, Marco Peer, Robert Sablatnig

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, accepted at ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15306 2025-10-20 cs.AI 79%

WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation

Kuang-Da Wang, Zhao Wang, Yotaro Shimose, Wei-Yao Wang, Shingo Takamatsu

机构 * National Yang Ming Chiao Tung University, Sony Group Corporation(National Yang Ming Chiao Tung University, Sony集团) Sony Group Corporation(Sony集团)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06979 2025-10-20 cs.CL 79%

Tackling Fake News in Bengali: Unraveling the Impact of Summarization vs. Augmentation on Pre-trained Language Models

Arman Sakif Chowdhury, G. M. Shahariar, Ahammed Tarik Aziz, Syed Mohibul Alam, Md. Azad Sheikh, Tanveer Ahmed Belal

机构 * Ahsanullah University of Science and Technology(阿沙努拉科学与技术大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

Comments Accepted, In Production

Journal ref SN COMPUT. SCI. 6, 917 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15842 2025-10-20 cs.CL cs.CV 77%

Paper2Web: Let's Make Your Paper Alive!

Yuhang Chen, Tianpeng Lv, Siyi Zhang, Yixiang Yin, Yao Wan, Philip S. Yu, Dongping Chen

机构 * ONE Lab, Huazhong University of Science and Technology(华中科技大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Maryland(马里兰大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments Under Review. Check https://github.com/YuhangChen1/Paper2All for the unified platform to streamline all academic presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15568 2025-10-20 cs.HC cs.AI 77%

The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems

Alexander Doudkin, Anton Voelker, Friedrich von Borries

机构 * Art of X UG (haftungsbeschraenkt)(Art of X 有限公司) HFBK Hamburg(汉堡艺术大学) Hamburg Open Online University (HOOU) program(汉堡开放在线大学(HOOU)计划)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments 10 pages, 2 figures, 2 tables. This project was collaboratively developed with the Art of X UG (haftungsbeschraenkt) AI Research team and HFBK Hamburg, with initial funding from the Hamburg Open Online University (HOOU) program

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11717 2025-10-20 cs.LG cs.AI cs.CL cs.CV 75%

WebInject: Prompt Injection Attack to Web Agents

Xilong Wang, John Bloch, Zedian Shao, Yuepeng Hu, Shuyan Zhou, Neil Zhenqiang Gong

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Appeared in EMNLP 2025 main conference. To better understand prompt injection attacks, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05482 2025-10-20 cs.GR cs.PL 75%

Imperative vs. Declarative Programming Paradigms for Open-Universe Scene Generation

Maxim Gumin, Do Heon Han, Seung Jean Yoo, Aditya Ganeshan, R. Kenny Jones, Rio Aguina-Kang, Stewart Morris, Daniel Ritchie

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15134 2025-10-20 cs.CL cs.AI 73%

FarsiMCQGen: a Persian Multiple-choice Question Generation Framework

Mohammad Heydari Rad, Rezvan Afari, Saeedeh Momtazi

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00022 2025-10-20 cs.CL cs.LG physics.ed-ph 73%

Scaling Physical Reasoning with the PHYSICS Dataset

Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Haonan He, Ning Ding, Yu Cheng, Shuyue Hu, Lei Bai, Dongzhan Zhou, Ganqu Cui, Peng Ye

机构 * Shanghai AI Laboratory(上海人工智能实验室) Harbin Institute of Technology(哈尔滨工业大学) Tsinghua University(清华大学) CUHK(香港中文大学) BUAA(北京航空航天大学) UESTC(电子科技大学) Soochow University(苏州大学) USTC(中国科学技术大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

Comments Accepted to the NeurIPS Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05649 2025-10-20 cs.IR cs.LG 70%

AI Guided Accelerator For Search Experience

Jayanth Yetukuri, Mehran Elyasi, Samarth Agrawal, Aritra Mandal, Rui Kong, Harish Vempati, Ishita Khan

机构 * eBay Inc(eBay公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.LG

Comments Accepted at SIGIR eCom'25. https://sigir-ecom.github.io/eCom25Papers/paper_25.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04995 2025-10-20 cs.HC cs.AI cs.DL 70%

Situated Epistemic Infrastructures: A Diagnostic Framework for Post-Coherence Knowledge

Matthew Kelly

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments 22 pages including references. Draft prepared for submission to Social Epistemology

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17669 2025-10-20 cs.CL 70%

Towards Human Cognition: Visual Context Guides Syntactic Priming in Fusion-Encoded Models

Bushi Xiao, Michael Bennie, Jayetri Bardhan, Daisy Zhe Wang

机构 * University of Florida(佛罗里达大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13990 2025-10-20 cs.SE 67%

RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust

Guangsheng Ou, Mingwei Liu, Yuxuan Chen, Yanlin Wang, Xin Peng, Zibin Zheng

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10444 2025-10-20 cs.CL cs.AI 62%

Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance

Jingyi Chen, Zhimeng Guo, Jiyun Chun, Pichao Wang, Andrew Perrault, Micha Elsner

机构 * Department of Linguistics, The Ohio State University, USA(语言学系,俄亥俄州立大学) Department of Computer Science and Engineering, The Ohio State University, USA(计算机科学与工程系,俄亥俄州立大学) Department of Information Sciences and Technology, Penn State University, USA(信息科学与技术系,宾夕法尼亚州立大学) Amazon, USA(亚马逊公司)

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15782 2025-10-20 cs.AI 57%

Demo: Guide-RAG: Evidence-Driven Corpus Curation for Retrieval-Augmented Generation in Long COVID

Philip DiGiacomo, Haoyang Wang, Jinrui Fang, Yan Leng, W Michael Brode, Ying Ding

机构 * Department of Computer Science University of Texas at Austin(计算机科学系得克萨斯大学奥斯汀分校) School of Information University of Texas at Austin(信息学院得克萨斯大学奥斯汀分校) McCombs School of Business University of Texas at Austin(麦库姆斯商学院得克萨斯大学奥斯汀分校) Dell Medical School University of Texas at Austin(德克萨斯大学奥斯汀分校戴尔医学学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments Accepted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: The Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15750 2025-10-20 cs.LG 57%

A Comprehensive Evaluation of Graph Neural Networks and Physics Informed Learning for Surrogate Modelling of Finite Element Analysis

Nayan Kumar Singh

专题命中 评测与基准 :pretraining(abstract);分类 cs.LG

Comments 14 pages, 6 figures, 5 tables. Code available at:https://github.com/SinghNayanKumar/DL-surrogate-modelling

详情

展开后加载摘要…

URL PDF HTML 收藏