arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8034 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8034 篇

2402.10908 2024-08-02 cs.CL cs.AI cs.HC cs.LG 67%

LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration

Hakan T. Otal, M. Abdullah Canbaz

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11190 2024-07-17 cs.CY cs.AI cs.CL 67%

In Silico Sociology: Forecasting COVID-19 Polarization with Large Language Models

Austin C. Kozlowski, Hyunku Kwon, James A. Evans

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11299 2024-07-16 cs.CV cs.AI cs.CL cs.LG 67%

SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant

Guohao Sun, Can Qin, Jiamian Wang, Zeyuan Chen, Ran Xu, Zhiqiang Tao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12963 2024-07-03 cs.RO cs.AI cs.CL cs.CV cs.LG 67%

AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents

Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Karol Hausman, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Sean Kirmani, Isabel Leal, Edward Lee, Sergey Levine, Yao Lu, Isabel Leal, Sharath Maddineni, Kanishka Rao, Dorsa Sadigh, Pannag Sanketi, Pierre Sermanet, Quan Vuong, Stefan Welker, Fei Xia, Ted Xiao, Peng Xu, Steve Xu, Zhuo Xu

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 26 pages, 9 figures, ICRA 2024 VLMNM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06102 2024-06-10 cs.CL cs.AI cs.LG 67%

Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, Mor Geva

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2024 (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17022 2024-06-05 cs.LG cs.AI cs.CL 67%

Controlled Decoding from Language Models

Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanping Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20213 2024-05-31 cs.AI cs.CL cs.LG 67%

PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization

Vijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas, Varre Chaitanya, Shwetha Somasundaram

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11070 2024-05-21 cs.AI cs.CL cs.LG 67%

Jill Watson: A Virtual Teaching Assistant powered by ChatGPT

Karan Taneja, Pratyusha Maiti, Sandeep Kakar, Pranav Guruprasad, Sanjeev Rao, Ashok K. Goel

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10745 2024-05-20 cs.LG cs.AI cs.CL 67%

Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings

Albert Sawczyn, Jakub Binkowski, Piotr Bielak, Tomasz Kajdanowicz

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted for LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09221 2024-05-16 cs.CL cs.AI cs.LG 67%

Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter

Josh McGiff, Nikola S. Nikolov

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 6 pages, Homophobia detection model available at: https://huggingface.co/JoshMcGiff/homophobiaBERT. The dataset used for this study is available at: https://huggingface.co/datasets/JoshMcGiff/HomophobiaDetectionTwitterX - This paper has been accepted by the 6th International Conference on Computing and Data Science (CONF-CDS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17999 2024-04-30 cs.CL cs.AI cs.LG 67%

MediFact at MEDIQA-CORR 2024: Why AI Needs a Human Touch

Nadia Saeed

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages, 4 figures, Clinical NLP 2024 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05875 2024-04-10 cs.CL cs.AI cs.LG 67%

CodecLM: Aligning Language Models with Tailored Synthetic Data

Zifeng Wang, Chun-Liang Li, Vincent Perot, Long T. Le, Jin Miao, Zizhao Zhang, Chen-Yu Lee, Tomas Pfister

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Findings of NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01358 2024-04-03 q-bio.QM cs.AI cs.CL cs.IR cs.LG cs.SI 67%

Utilizing AI and Social Media Analytics to Discover Adverse Side Effects of GLP-1 Receptor Agonists

Alon Bartal, Kathleen M. Jagodnik, Nava Pliskin, Abraham Seidmann

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 19 pages, 7 figures, 3 tables, 1 Appendix table

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01786 2024-03-29 cs.AI cs.CL cs.HC cs.LG 67%

COA-GPT: Generative Pre-trained Transformers for Accelerated Course of Action Development in Military Operations

Vinicius G. Goecks, Nicholas Waytowich

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at the NATO Science and Technology Organization Symposium (ICMCIS) organized by the Information Systems Technology (IST) Panel, IST-205-RSY - the ICMCIS, held in Koblenz, Germany, 23-24 April 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02151 2024-03-28 cs.CL cs.AI cs.LG 67%

Identifying the Correlation Between Language Distance and Cross-Lingual Transfer in a Multilingual Representation Space

Fred Philippy, Siwen Guo, Shohreh Haddadan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments SIGTYP Workshop 2023 (co-located with EACL 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05224 2024-03-26 cs.CV cs.AI cs.CL cs.LG 67%

Do Vision and Language Encoders Represent the World Similarly?

Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Mohamed El Amine Seddik, Karttikeya Mangalam, Noel E. O'Connor

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11490 2024-03-19 cs.CV cs.AI cs.CL cs.LG 67%

LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation

Suhyeon Lee, Won Jun Kim, Jinho Chang, Jong Chul Ye

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 21 pages, 8 figures; ICLR 2024 (poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04795 2024-03-11 cs.CL cs.AI cs.LG 67%

Large Language Models in Fire Engineering: An Examination of Technical Questions Against Domain Knowledge

Haley Hostetter, M. Z. Naser, Xinyan Huang, John Gales

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02325 2024-03-05 cs.CV cs.AI cs.CL cs.LG 67%

Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training

David Wan, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Project website: https://contrastive-region-guidance.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11259 2024-03-05 cs.LG cs.AI cs.CE cs.CL 67%

Domain-Agnostic Molecular Generation with Chemical Feedback

Yin Fang, Ningyu Zhang, Zhuo Chen, Lingbing Guo, Xiaohui Fan, Huajun Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06668 2024-02-15 cs.LG cs.AI cs.CL 67%

In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Sheng Liu, Haotian Ye, Lei Xing, James Zou

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17179 2024-02-12 cs.LG cs.AI cs.CL 67%

Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Xidong Feng, Ziyu Wan, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, Jun Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11562 2024-01-26 cs.AI cs.CL cs.CV cs.LG 67%

A Survey of Reasoning with Foundation Models

Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Ding, Hongyang Li, Mengzhe Geng, Yue Wu, Wenhai Wang, Junsong Chen, Zhangyue Yin, Xiaozhe Ren, Jie Fu, Junxian He, Wu Yuan, Qi Liu, Xihui Liu, Yu Li, Hao Dong, Yu Cheng, Ming Zhang, Pheng Ann Heng, Jifeng Dai, Ping Luo, Jingdong Wang, Ji-Rong Wen, Xipeng Qiu, Yike Guo, Hui Xiong, Qun Liu, Zhenguo Li

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 20 Figures, 160 Pages, 750+ References, Project Page https://github.com/reasoning-survey/Awesome-Reasoning-Foundation-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09454 2024-01-19 cs.CV cs.AI cs.CL cs.LG 67%

Voila-A: Aligning Vision-Language Models with User's Gaze Attention

Kun Yan, Lei Ji, Zeyu Wang, Yuntao Wang, Nan Duan, Shuai Ma

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19680 2024-01-17 cs.CL cs.AI cs.LG 67%

Integrating Pre-trained Language Model into Neural Machine Translation

Soon-Jae Hwang, Chang-Sung Jeong

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11671 2024-01-05 cs.CL cs.AI cs.LG 67%

Evaluating Language-Model Agents on Realistic Autonomous Tasks

Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R. Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, Paul Christiano

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05503 2023-12-12 cs.CL cs.AI cs.LG 67%

Aligner: One Global Token is Worth Millions of Parameters When Aligning Large Language Models

Zhou Ziheng, Yingnian Wu, Song-Chun Zhu, Demetri Terzopoulos

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 81 pages, 77 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10313 2023-12-06 cs.CL cs.AI cs.LG 67%

Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, Yi Ma

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10093 2023-11-08 cs.CV cs.AI cs.CL cs.LG 67%

Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and Detection

Kyle Buettner, Adriana Kovashka

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at Winter Conference on Applications of Computer Vision (WACV), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02105 2023-11-07 cs.LG cs.AI cs.CY 67%

Making Harmful Behaviors Unlearnable for Large Language Models

Xin Zhou, Yi Lu, Ruotian Ma, Tao Gui, Qi Zhang, Xuanjing Huang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments work in process

详情

展开后加载摘要…

URL PDF HTML 收藏