arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2403.08295 2024-04-17 cs.CL cs.AI 62%

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu-hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, Kathleen Kenealy

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09574 2024-04-16 cs.LG cs.AI 62%

Predicting and Analyzing Pedestrian Crossing Behavior at Unsignalized Crossings

Chi Zhang, Janis Sprenger, Zhongjun Ni, Christian Berger

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 10 figures, 4 tables. Accepted in 2024 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07484 2024-04-16 cs.CL cs.AI 62%

Psychometric Predictive Power of Large Language Models

Tatsuki Kuribayashi, Yohei Oseki, Timothy Baldwin

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 23 pages; Findings of NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06971 2024-04-11 cs.CV cs.AI cs.LG 62%

TrajPRed: Trajectory Prediction with Region-based Relation Learning

Chen Zhou, Ghassan AlRegib, Armin Parchami, Kunjan Singh

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02736 2024-04-11 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

Discovering Closed-Loop Failures of Vision-Based Controllers via Reachability Analysis

Kaustav Chakraborty, Somil Bansal

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Robotics and Automation Letters 8.5 (2023): 2692-2699

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12508 2024-04-05 cs.LG cs.AI 62%

SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation

Chongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong, Dennis Wei, Sijia Liu

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by ICLR 2024 as a Spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16611 2024-04-03 cs.CL cs.AI cs.HC 62%

Understanding the Dataset Practitioners Behind Large Language Model Development

Crystal Qian, Emily Reif, Minsuk Kahng

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 2 figures. To be published in In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA '24). Revised to reflect updates from CHI LBW reviewer feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00983 2024-04-02 cs.LG cs.AI 62%

Continual Learning for Smart City: A Survey

Li Yang, Zhipeng Luo, Shiming Zhang, Fei Teng, Tianrui Li

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint. Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04357 2024-03-28 cs.CL cs.AI 62%

Dial-MAE: ConTextual Masked Auto-Encoder for Retrieval-based Dialogue Systems

Zhenpeng Su, Xing Wu, Wei Zhou, Guangyuan Ma, Songlin Hu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13257 2024-03-27 cs.CL cs.AI 62%

Visual Grounding Helps Learn Word Meanings in Low-Data Regimes

Chengxu Zhuang, Evelina Fedorenko, Jacob Andreas

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13830 2024-03-22 q-bio.BM cs.CL cs.LG 62%

Bridging Text and Molecule: A Survey on Multimodal Frameworks for Molecule

Yi Xiao, Xiangxin Zhou, Qiang Liu, Liang Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13514 2024-03-21 cs.CL cs.CY 62%

How Gender Interacts with Political Values: A Case Study on Czech BERT Models

Adnan Al Ali, Jindřich Libovický

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments 11 pages, 2 figures; LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09795 2024-03-18 cs.CR cs.AI cs.CL 62%

Helpful or Harmful? Exploring the Efficacy of Large Language Models for Online Grooming Prevention

Ellie Prosser, Matthew Edwards

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07699 2024-03-15 cs.CV cs.AI cs.LG 62%

VeCLIP: Improving CLIP Training via Visual-enriched Captions

Zhengfeng Lai, Haotian Zhang, Bowen Zhang, Wentao Wu, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang, Meng Cao

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments CV/ML

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06914 2024-03-13 cs.CL cs.AI 62%

MEND: Meta dEmonstratioN Distillation for Efficient and Effective In-Context Learning

Yichuan Li, Xiyao Ma, Sixing Lu, Kyumin Lee, Xiaohu Liu, Chenlei Guo

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12021 2024-03-12 cs.CL cs.AI 62%

Synergistic Anchored Contrastive Pre-training for Few-Shot Relation Extraction

Da Luo, Yanglei Gan, Rui Hou, Run Lin, Qiao Liu, Yuxiang Cai, Wannian Gao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12151 2024-03-05 cs.CL cs.AI 62%

Transformer-based Causal Language Models Perform Clustering

Xinbo Wu, Lav R. Varshney

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Added new experimental results and fixed some errors

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18807 2024-03-01 cs.CL cs.AI 62%

On the Decision-Making Abilities in Role-Playing using Large Language Models

Chenglei Shen, Guofu Xie, Xiao Zhang, Jun Xu

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18096 2024-02-29 cs.LG cs.AI 62%

No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

June Yong Yang, Byeongwook Kim, Jeongin Bae, Beomseok Kwon, Gunho Park, Eunho Yang, Se Jung Kwon, Dongsoo Lee

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16305 2024-02-27 cs.LG cs.AI 62%

Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion

Xuantong Liu, Tianyang Hu, Wenjia Wang, Kenji Kawaguchi, Yuan Yao

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09443 2024-02-16 eess.SP cs.AI cs.LG 62%

Review of algorithms for predicting fatigue using EEG

Ildar Rakhmatulin

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2401.15766

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08088 2024-02-14 cs.AI cs.LG eess.IV 62%

Out-of-Distribution Detection and Data Drift Monitoring using Statistical Process Control

Ghada Zamzmi, Kesavan Venkatesh, Brandon Nelson, Smriti Prathapan, Paul H. Yi, Berkman Sahiner, Jana G. Delfino

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07344 2024-02-13 cs.LG cs.AI 62%

Measurement Scheduling for ICU Patients with Offline Reinforcement Learning

Zongliang Ji, Anna Goldenberg, Rahul G. Krishnan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Extended Abstract presented at Machine Learning for Health (ML4H) symposium 2023, December 10th, 2023, New Orleans, United States, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06185 2024-02-12 cs.CV cs.AI cs.LG 62%

Development and validation of an artificial intelligence model to accurately predict spinopelvic parameters

Edward S. Harake, Joseph R. Linzey, Cheng Jiang, Rushikesh S. Joshi, Mark M. Zaki, Jaes C. Jones, Siri S. Khalsa, John H. Lee, Zachary Wilseck, Jacob R. Joseph, Todd C. Hollon, Paul Park

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 5 figures, to appear in Journal of Neurosurgery: Spine

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05627 2024-02-09 cs.LG cs.AI cs.CV q-bio.NC 62%

Binding Dynamics in Rotating Features

Sindy Löwe, Francesco Locatello, Max Welling

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04232 2024-02-08 cs.AI cs.CL 62%

Can Generative Agents Predict Emotion?

Ciaran Regan, Nanami Iwahashi, Shogo Tanaka, Mizuki Oka

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16123 2024-02-08 cs.HC cs.AI cs.CV cs.LG 62%

Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual Drivers

Amr Gomaa, Guillermo Reyes, Michael Feld, Antonio Krüger

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in the Proceedings of the 29th International Conference on Intelligent User Interfaces (IUI'24), March 18--21, 2024, in Greenville, SC, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01516 2024-02-07 cs.CV cs.AI cs.LG cs.MM 62%

MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval

Zijun Long, George Killick, Richard McCreadie, Gerardo Aragon Camarasa

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03014 2024-02-06 cs.LG cs.AI 62%

Whom to Trust? Elective Learning for Distributed Gaussian Process Regression

Zewen Yang, Xiaobing Dai, Akshat Dubey, Sandra Hirche, Georges Hattab

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, conference preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10246 2024-02-06 cs.LG cs.AI stat.ML 62%

Surprisal Driven $k$-NN for Robust and Interpretable Nonparametric Learning

Amartya Banerjee, Christopher J. Hazard, Jacob Beel, Cade Mack, Jack Xia, Michael Resnick, Will Goddin

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏