arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2311.09809 2023-11-17 cs.LO cs.AI cs.LG 62%

Comparing Differentiable Logics for Learning Systems: A Research Preview

Thomas Flinkow, Barak A. Pearlmutter, Rosemary Monahan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments In Proceedings FMAS 2023, arXiv:2311.08987

Journal ref EPTCS 395, 2023, pp. 17-29

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08993 2023-11-16 cs.CL cs.AI 62%

When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks

Hao Peng, Xiaozhi Wang, Jianhui Chen, Weikai Li, Yunjia Qi, Zimu Wang, Zhili Wu, Kaisheng Zeng, Bin Xu, Lei Hou, Juanzi Li

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07583 2023-11-15 cs.CL cs.AI 62%

Cross-Dialect Sentence Transformation: A Comparative Analysis of Language Models for Adapting Sentences to British English

Shruti Dutta, Shashwat Mookherjee

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06650 2023-11-14 cs.LG cs.AI cs.NE cs.RO stat.ML 62%

Interpreting Neural Policies with Disentangled Tree Representations

Tsun-Hsuan Wang, Wei Xiao, Tim Seyde, Ramin Hasani, Daniela Rus

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01150 2023-11-03 cs.CL cs.AI 62%

Revisiting the Knowledge Injection Frameworks

Peng Fu, Yiming Zhang, Haobo Wang, Weikang Qiu, Junbo Zhao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP 2023 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00235 2023-11-02 stat.ML cs.AI cs.LG 62%

Implicit biases in multitask and continual learning from a backward error analysis perspective

Benoit Dherin

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted in Mathematics of Modern Machine Learning Workshop at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17956 2023-11-02 cs.CV cs.AI cs.CL 62%

Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare

Junling Liu, Ziming Wang, Qichen Ye, Dading Chong, Peilin Zhou, Yining Hua

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19961 2023-11-01 cs.LG cs.AI 62%

ExPT: Synthetic Pretraining for Few-Shot Experimental Design

Tung Nguyen, Sudhanshu Agrawal, Aditya Grover

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 2023 Conference on Neural Information Processing Systems (NeurIPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19957 2023-11-01 cs.LG cs.AI cs.CV 62%

Deep Learning for Spatiotemporal Big Data: A Vision on Opportunities and Challenges

Zhe Jiang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16999 2023-10-31 cs.CV cs.AI cs.LG 62%

Three Towers: Flexible Contrastive Learning with Pretrained Image Models

Jannik Kossen, Mark Collier, Basil Mustafa, Xiao Wang, Xiaohua Zhai, Lucas Beyer, Andreas Steiner, Jesse Berent, Rodolphe Jenatton, Efi Kokiopoulou

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04717 2023-10-31 cs.LG cs.AI 62%

On the Sensitivity of Reward Inference to Misspecified Human Models

Joey Hong, Kush Bhatia, Anca Dragan

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments published as a paper in ICLR 2023; 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13788 2023-10-30 cs.CL cs.AI 62%

Can Large Language Models Capture Dissenting Human Voices?

Noah Lee, Na Min An, James Thorne

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments To appear at EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15080 2023-10-27 cs.CL cs.AI 62%

Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

Geewook Kim, Hodong Lee, Daehee Kim, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, Seunghyun Park

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 22 pages; To appear at EMNLP 2023 Main Conference (Project page: https://naver-ai.github.io/cream )

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15188 2023-10-25 cs.LG cond-mat.mtrl-sci cs.AI 62%

Deep Learning Approaches for Dynamic Mechanical Analysis of Viscoelastic Fiber Composites

Victor Hoffmann, Ilias Nahmed, Parisa Rastin, Guénaël Cabanes, Julien Boisse

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 5 figures, https://hal.science/hal-04250557

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14799 2023-10-24 cs.CL cs.AI 62%

Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages

Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, Wanxiang Che

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP2023 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09343 2023-10-24 cs.CL cs.AI 62%

Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents

Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, Jinyoung Yeo

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 25 pages, 8 figures, Accepted to EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13596 2023-10-23 cs.CL cs.AI 62%

MarineGPT: Unlocking Secrets of Ocean to the Public

Ziqiang Zheng, Jipeng Zhang, Tuan-Anh Vu, Shizhe Diao, Yue Him Wong Tim, Sai-Kit Yeung

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments work in progress. Code and data will be available at https://github.com/hkust-vgd/MarineGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07328 2023-10-23 cs.CL cs.AI 62%

An Empirical Study of Instruction-tuning Large Language Models in Chinese

Qingyi Si, Tong Wang, Zheng Lin, Xu Zhang, Yanan Cao, Weiping Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11084 2023-10-20 cs.LG cs.CL cs.CV 62%

Test-Time Distribution Normalization for Contrastively Learned Vision-language Models

Yifei Zhou, Juntao Ren, Fengyu Li, Ramin Zabih, Ser-Nam Lim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2023, project webpage: https://fengyuli-dev.github.io/dn-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09344 2023-10-17 cs.LG cs.AI cs.CV 62%

Beyond Distribution Shift: Spurious Features Through the Lens of Training Dynamics

Nihal Murali, Aahlad Puli, Ke Yu, Rajesh Ranganath, Kayhan Batmanghelich

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Main paper: 12 pages, 2 tables, and 10 figures. Supplementary: 10 pages and 9 figures. Accepted in TMLR23 (https://openreview.net/pdf?id=Tkvmt9nDmB)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06200 2023-10-12 cs.CL cs.LG 62%

The Importance of Prompt Tuning for Automated Neuron Explanations

Justin Lee, Tuomas Oikarinen, Arjun Chatha, Keng-Chi Chang, Yilan Chen, Tsui-Wei Weng

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06702 2023-10-11 cs.CL cs.LG cs.SD eess.AS 62%

Temporally Aligning Long Audio Interviews with Questions: A Case Study in Multimodal Data Integration

Piyush Singh Pasi, Karthikeya Battepati, Preethi Jyothi, Ganesh Ramakrishnan, Tanmay Mahapatra, Manoj Singh

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Work Accepted in IJCAI-23- AI and Social Good Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03666 2023-10-06 cs.CL cs.AI 62%

MapperGPT: Large Language Models for Linking and Mapping Entities

Nicolas Matentzoglu, J. Harry Caufield, Harshad B. Hegde, Justin T. Reese, Sierra Moxon, Hyeongsik Kim, Nomi L. Harris, Melissa A Haendel, Christopher J. Mungall

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16145 2023-09-29 cs.CL cs.CY cs.HC 62%

The Confidence-Competence Gap in Large Language Models: A Cognitive Study

Aniket Kumar Singh, Suman Devkota, Bishal Lamichhane, Uttam Dhakal, Chandra Dhakal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments 19 pages, 8 Figures, to be published in a journal (Journal TBD), All Authors contributed equally and were Supervised by Chandra Dhakal

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13483 2023-09-26 stat.ML cs.AI cs.LG 62%

Enhancing Prediction and Analysis of UK Road Traffic Accident Severity Using AI: Integration of Machine Learning, Econometric Techniques, and Time Series Forecasting in Public Health Research

Md Abu Sufian, Jayasree Varadarajan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 36

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16342 2023-09-26 cs.CV cs.AI cs.CL 62%

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

Reuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer, Justin Salamon, Oriol Nieto, Bryan Russell, Kate Saenko

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11196 2023-09-21 cs.LG cs.AI cs.CR cs.SC 62%

When to Trust AI: Advances and Challenges for Certification of Neural Networks

Marta Kwiatkowska, Xiyue Zhang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08773 2023-09-19 cs.SD cs.AI cs.LG eess.AS 62%

Enhance audio generation controllability through representation similarity regularization

Yangyang Shi, Gael Le Lan, Varun Nagaraja, Zhaoheng Ni, Xinhao Mei, Ernie Chang, Forrest Iandola, Yang Liu, Vikas Chandra

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03592 2023-09-15 cs.CL cs.AI q-bio.NC 62%

Testing the limits of natural language models for predicting human language judgments

Tal Golan, Matthew Siegelman, Nikolaus Kriegeskorte, Christopher Baldassano

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref Nature Machine Intelligence (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03905 2023-09-13 cs.MM cs.CL cs.CV cs.LG cs.SD eess.AS 62%

ImageBind-LLM: Multi-modality Instruction Tuning

Jiaming Han, Renrui Zhang, Wenqi Shao, Peng Gao, Peng Xu, Han Xiao, Kaipeng Zhang, Chris Liu, Song Wen, Ziyu Guo, Xudong Lu, Shuai Ren, Yafei Wen, Xiaoxin Chen, Xiangyu Yue, Hongsheng Li, Yu Qiao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Code is available at https://github.com/OpenGVLab/LLaMA-Adapter

详情

展开后加载摘要…

URL PDF HTML 收藏