Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
A Scalable Framework for Evaluating Health Language Models
可扩展的健康语言模型评估框架
Neil Mallinar, A. Ali Heydari, Xin Liu, Anthony Z. Faranesh, Brent Winslow, Nova Hammerquist, Benjamin Graef, Cathy Speed, Mark Malhotra, Shwetak Patel, Javier L. Prieto, Daniel McDuff, Ahmed A. Metwally
机构
*
Google Research(谷歌研究)
专题命中
评测与基准
:language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI
机构
*
organization= School of Computer Science
;
Technology, Dalian University of Technology , addressline= No.2 Linggong Road, Ganjingzi District , city= Dalian , postcode= 116024 , country= China
;
organization= School of Information Engineering, Liaodong University , addressline= No.116 Linjiang Back Street, Zhenan District , city= Dandong , postcode= 118001 , country= China
;
organization= School of Information Engineering, Dalian Ocean University , addressline= No. 2-52, Heishijiao Street, Shahekou District , city= Dalian , postcode= 116023 , country= China
;
organization= Information Technology Center, Qinghai University , addressline= 251 Ningda Road, Chengbei District , city= Xining , postcode= 810016 , country= China
;
organization= School of Electronic
;
Information Engineering, Liaoning Technical University , addressline= 188 Longwan South Street, Sijiatun District , city= Huludao , postcode= 125105 , country= China
;
organization= School of Artificial Intelligence, Tianjin Normal University , addressline= 393 Binshui West Road, Xiqing District , city= Tianjin , postcode= 300387 , country= China
;
organization= Tencent (Dalian Northern Interactive Entertainment Technology Co., Ltd.) , addressline= 21/F, Tencent Building, No. 26 Jingxian St, Ganjingzi District , city= Dalian , postcode= 116085 , country= China
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI Agents for Inventory Control: Human-LLM-OR Complementarity
面向库存控制的AI代理:人类-大语言模型-运筹学互补性
Jackie Baek, Yaopeng Fu, Will Ma, Tianyi Peng
机构
*
Stern School of Business, New York University(纽约大学斯特恩商学院)
;
Columbia University(哥伦比亚大学)
;
Graduate School of Business and Data Science Institute, Columbia University(哥伦比亚大学商学院与数据科学研究院)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
重新定义评估标准:一种统一评估韩语模型能力的框架
Hanwool Lee, Dasol Choi, Sooyong Kim, Ilgyun Jeong, Sangwon Baek, Guijin Son, Inseon Hwang, Naeun Lee, Seunghyeok Hong
机构
*
AIM Intelligence
;
Seoul National University(首尔国立大学)
;
A.I.MATICS
;
TigerCompany
;
Catius
;
National Assembly of Korea(韩国国会)
;
Coupang
;
Hankuk University of Foreign Studies(韩国外交大学)
专题命中
评测与基准
:language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI
From Instruction to Output: The Role of Prompting in Modern NLG
从指令到输出:提示在现代自然语言生成中的作用
Munazza Zaib, Elaf Alhazmi
机构
*
Faculty of Information Technology, Monash University(莫纳什大学信息科技学院)
;
School of Computer, Data and Mathematical Sciences, Western Sydney University(西澳大学计算机、数据与数学科学学院)
;
School of Computing, Macquarie University(麦考瑞大学计算学院)
专题命中
评测与基准
:prompting(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
解析LLM作为裁判:用于推理时间扩展的可分析模型
Indranil Halder, Cengiz Pehlevan
机构
*
John A. Paulson School of Engineering And Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)
;
Center for Brain Science, Harvard University(哈佛大学脑科学中心)
;
Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Don't Always Pick the Highest-Performing Model: An Information Theoretic View of LLM Ensemble Selection
不要总是选择表现最好的模型:对大语言模型集成选择的信息论视角
Yigit Turkmen, Baturalp Buyukates, Melih Bastopcu
机构
*
Department of Electrical and Electronics Engineering, Bilkent University, Ankara, Turkey(电气与电子工程系,比尔肯大学,安卡拉,土耳其)
;
School of Computer Science, University of Birmingham, Birmingham, UK(计算机科学学院,伯明翰大学,伯明翰,英国)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
CommentsPart of this work (RQ1) has been published at the 2026 IEEE/ACM 48th International Conference on Software Engineering (ICSE-SEIP 2026), DOI: 10.1145/3786583.3786904. The published version is also available on arXiv at arXiv:2602.04449
机构
*
Dept. of Computer Science Stevens Institute of Technology(计算机科学系 斯坦福理工学院)
;
Dept. of Computer Science Binghamton University(计算机科学系 哈伯里顿大学)
;
Engineering The Ohio State University(工程学 俄亥俄州立大学)
;
Learning Division Argonne National Laboratory(学习部 阿贡国家实验室)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
TextMineX: 用于人道主义排雷行动的数据集、评估框架及基于本体的LLM流水线
Chenyue Zhou, Gürkan Solmaz, Flavio Cirillo, Kiril Gashteovski, Jonathan Fürst
机构
*
NEC Laboratories Europe(NEC欧洲实验室)
;
University of Stuttgart(斯图加特大学)
;
VAGO Solutions(VAGO解决方案)
;
Zurich University of Applied Sciences(苏黎世应用科学大学)
;
CAIR, Ss. Cyril and Methodius University of Skopje, North Macedonia(CAIR,斯科普耶塞尔维亚·梅托迪乌斯大学,北马其顿)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI