Capability-Based Scaling Trends for LLM-Based Red-Teaming
基于能力的LLM红队测试规模趋势
Alexander Panfilov, Paul Kassianik, Maksym Andriushchenko, Jonas Geiping
机构
*
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Tübingen AI Center(图宾根人工智能中心)
;
Foundation AI – Cisco Systems Inc.(AI基础研究机构——思科系统公司)
;
EPFL(苏黎世联邦理工学院)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Don't Always Pick the Highest-Performing Model: An Information Theoretic View of LLM Ensemble Selection
不要总是选择表现最好的模型:对大语言模型集成选择的信息论视角
Yigit Turkmen, Baturalp Buyukates, Melih Bastopcu
机构
*
Department of Electrical and Electronics Engineering, Bilkent University, Ankara, Turkey(电气与电子工程系,比尔肯大学,安卡拉,土耳其)
;
School of Computer Science, University of Birmingham, Birmingham, UK(计算机科学学院,伯明翰大学,伯明翰,英国)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
G-LNS: Generative Large Neighborhood Search for LLM-Based Automatic Heuristic Design
G-LNS:基于生成式大邻域搜索的LLM自动启发式设计
Baoyun Zhao, He Wang, Liang Zeng
机构
*
Software College, Northeastern University(东北大学软件学院)
;
International Centre for Theoretical Physics Asia-Pacific, University of Chinese Academy of Sciences(中国科学院大学国际理论物理亚太中心)
;
Taiji Laboratory for Gravitational Wave Universe, University of Chinese Academy of Sciences(中国科学院大学太极引力波宇宙实验室)
;
Tsinghua University(清华大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
Game-Theoretic Co-Evolution for LLM-Based Heuristic Discovery
基于博弈论的LLM启发式发现共演化
Xinyi Ke, Kai Li, Junliang Xing, Yifan Zhang, Jian Cheng
机构
*
C2DL, Institute of Automation, Chinese Academy of Sciences(C2DL,自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
从多模态数据集构建通用基础模型用于三维计算机断层扫描
Ibrahim Ethem Hamamci, Sezgin Er, Chenyu Wang, Furkan Almas, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Irem Dogan, Omer Faruk Durugol, Benjamin Hou, Suprosanna Shit, Weicheng Dai, Murong Xu, Hadrien Reynaud, Muhammed Furkan Dasdelen, Bastian Wittmann, Tamaz Amiranashvili, Enis Simsar, Mehmet Simsar, Emine Bensu Erdemir, Abdullah Alanbay, Anjany Sekuboyina, Berkan Lafci, Ahmet Kaplan, Zhiyong Lu, Malgorzata Polacin, Bernhard Kainz, Christian Bluethgen, Kayhan Batmanghelich, Mehmet Kemal Ozdemir, Bjoern Menze
机构
*
Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系)
;
ETH AI Center, ETH Zurich(苏黎世联邦理工学院人工智能中心)
;
International School of Medicine, Istanbul Medipol University(伊斯坦布尔梅迪波尔大学国际医学院)
;
Department of Electrical and Computer Engineering, Boston University(波士顿大学电气与计算机工程系)
;
Division of Intramural Research, National Institutes of Health(美国国立卫生研究院内部研究部)
;
Department of Computing, Imperial College London(伦敦帝国理工学院计算机系)
;
Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系)
;
Institute for Diagnostic and Interventional Radiology, University Hospital Zurich(苏黎世大学医院诊断与介入放射学研究所)
;
Department Artificial Intelligence in Biomedical Engineering, FAU Erlangen-Nürnberg(埃尔朗根-纽伦堡工业大学生物医学工程人工智能系)
专题命中
评测与基准
:foundation model(title);large language model(abstract);language model(abstract);pretraining(abstract)
机构
*
Department of Computer and Network Engineering, United Arab Emirates University, UAE(计算机与网络工程系,阿拉伯联合酋长国大学)
;
G Research Center (6GRC), Khalifa University, UAE(6G研究中心(6GRC),哈利法大学)
Shadman Rabby, Md. Hefzul Hossain Papon, Sabbir Ahmed, Nokimul Hasan Arif, A. B. M. Ashikur Rahman, Irfan Ahmad
机构
*
University of Dhaka(达卡大学)
;
Daffodil International University(达福迪国际大学)
;
Islamic University of Technology(伊斯兰技术大学)
;
University of Central Florida(佛罗里达中央大学)
;
King Fahad University of Petroleum and Minerals(国王法赫德石油与矿物大学)
;
SDAIA - KFUPM Joint research Center for Artificial Intelligence(SDAIA-KFUPM人工智能联合研究中心)
Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
用恒星进行推理:一个用于代理科学推理的太阳物理数据集和基准
Kevin Lee, Russell Spiewak, James Walsh
机构
*
Frontier Development Lab(前沿发展实验室)
;
Department of Mechanical and Aerospace Engineering, UCLA(机械与航空航天工程系,加州大学洛杉矶分校)
;
Trillium Technologies Inc.(Trillium技术公司)
;
Department of Engineering, University of Cambridge(工程系,剑桥大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI、cs.LG