arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8016 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8016 篇

2407.05320 2024-07-09 cs.AI 79%

KAE: A Property-based Method for Knowledge Graph Alignment and Extension

Daqian Shi, Xiaoyue Li, Fausto Giunchiglia

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2405.02463

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05054 2024-07-09 cs.CL 79%

Cross-Lingual Word Alignment for ASEAN Languages with Contrastive Learning

Jingshen Zhang, Xinying Qiu, Teng Shen, Wenyu Wang, Kailin Zhang, Wenhe Feng

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01646 2024-07-03 cs.SE cs.AI 79%

ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization

Chunrong Fang, Weisong Sun, Yuchen Chen, Xiao Chen, Zhao Wei, Quanjun Zhang, Yudu You, Bin Luo, Yang Liu, Zhenyu Chen

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to IEEE Transactions on Software Engineering (TSE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19255 2024-06-28 cs.CV cs.CL 79%

Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment

Hao Fei, Shengqiong Wu, Meishan Zhang, Min Zhang, Tat-Seng Chua, Shuicheng Yan

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted by IEEE TPAMI 2024

Journal ref [J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19032 2024-06-28 cs.CL 79%

Improving Weak-to-Strong Generalization with Reliability-Aware Alignment

Yue Guo, Yi Yang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13561 2024-06-27 cs.CL cs.CV 79%

Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment

Yunxin Li, Xinyu Chen, Baotian Hu, Haoyuan Shi, Min Zhang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 12 pages,4 figures; Accepted by ACL 2024 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13894 2024-06-21 cs.CV cs.CY 79%

Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events

Mohammad Abu Tami, Huthaifa I. Ashqar, Mohammed Elhenawy

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13357 2024-06-21 cs.CL cs.SD eess.AS 79%

Transferable speech-to-text large language model alignment module

Boyong Wu, Chao Yan, Haoran Pu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted by InterSpeech 2024; 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05678 2024-06-21 cs.HC cs.CL 79%

Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment

Yoonsu Kim, Kihoon Son, Seoyoung Kim, Juho Kim

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06228 2024-06-12 cs.CL 79%

Understanding Cross-Lingual Alignment -- A Survey

Katharina Hämmerl, Jindřich Libovický, Alexander Fraser

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Camera-ready version, ACL Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06049 2024-06-11 cs.CY 79%

Enhancing Food Safety in Supply Chains: The Potential Role of Large Language Models in Preventing Campylobacter Contamination

Asaf Tzachor

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

Comments 29 pages, 1 figure, 3 boxes

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05505 2024-06-11 cs.IR cs.AI 79%

I-SIRch: AI-Powered Concept Annotation Tool For Equitable Extraction And Analysis Of Safety Insights From Maternity Investigations

Mohit Kumar Singh, Georgina Cosma, Patrick Waterson, Jonathan Back, Gyuchan Thomas Jun

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04799 2024-06-10 cs.CL 79%

Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages

Shih-Cheng Huang, Pin-Zu Li, Yu-Chi Hsu, Kuang-Ming Chen, Yu Tung Lin, Shih-Kai Hsiao, Richard Tzong-Han Tsai, Hung-yi Lee

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments ACL 2024 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02915 2024-06-06 cs.CV cs.LG 79%

Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models

Jinhao Li, Haopeng Li, Sarah Erfani, Lei Feng, James Bailey, Feng Liu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments 22 pages, 16 figures, published to ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00117 2024-05-29 cs.CL 79%

BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B

Pranav Gade, Simon Lermen, Charlie Rogers-Smith, Jeffrey Ladish

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00916 2024-05-29 cs.CL cs.SD eess.AS 79%

BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Chen Wang, Minpeng Liao, Zhongqiang Huang, Jinliang Lu, Junhong Wu, Yuchen Liu, Chengqing Zong, Jiajun Zhang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15787 2024-05-28 cs.IR cs.CL 79%

Extracting chemical food safety hazards from the scientific literature automatically using large language models

Neris Özen, Wenjuan Mu, Esther D. van Asselt, Leonieke M. van den Bulk

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15430 2024-05-27 cs.LG cs.LO 79%

Counterexample-Guided Repair of Reinforcement Learning Systems Using Safety Critics

David Boetius, Stefan Leue

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments 7 pages + references

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15202 2024-05-27 cs.CL cs.CR 79%

Cross-Task Defense: Instruction-Tuning LLMs for Content Safety

Yu Fu, Wen Xiao, Jia Chen, Jiachen Li, Evangelos Papalexakis, Aichi Chien, Yue Dong

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments accepted to NAACL2024 TrustNLP workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07162 2024-05-17 cs.RO cs.AI 79%

Learning Reward for Robot Skills Using Large Language Models via Self-Alignment

Yuwei Zeng, Yao Mu, Lin Shao

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08154 2024-05-15 cs.HC cs.AI 79%

LLM Theory of Mind and Alignment: Opportunities and Risks

Winnie Street

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Journal ref Proceedings of Workshop on Theory of Mind in Human-AI Interaction at CHI 2024 (ToMinHAI at CHI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03939 2024-05-08 cs.CL 79%

Long Context Alignment with Short Instructions and Synthesized Positions

Wenhao Wu, Yizhong Wang, Yao Fu, Xiang Yue, Dawei Zhu, Sujian Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments preview

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03699 2024-05-08 cs.HC cs.AI 79%

HCC Is All You Need: Alignment-The Sensible Kind Anyway-Is Just Human-Centered Computing

Eric Gilbert

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00462 2024-05-06 cs.LG cs.RO 79%

Zero-shot Safety Prediction for Autonomous Robots with Foundation World Models

Zhenjiang Mao, Siqi Dai, Yuang Geng, Ivan Ruchkin

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments Presented at the Back to the Future-Robot Learning Going Probabilistic Workshop, co-located with ICRA 2024. https://openreview.net/forum?id=gHhBNIq9Cs

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08774 2024-04-30 physics.soc-ph cs.AI 79%

Discussion of Loop Expansion and Introduction of Series Cutting Functions to Local Potential Approximation: Complexity Analysis Using Green's Functions, Cutting Of Nth-Order Social Interactions For Progressive Safety

Yasuko Kawahata

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments In this study, we focus on the aforementioned paper, "Examination Kubo-Matsubara Green's Function Of The Edwards-Anderson Model: Extreme Value Information Flow Of Nth-Order Interpolated Extrapolation Of Zero Phenomena Using The Replica Method (2024)"

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05183 2024-04-09 cs.CV cs.LG 79%

Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset

Chih-Chung Hsu, Chia-Ming Lee, Chun-Hung Sun, Kuang-Ming Wu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments MULA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04659 2024-04-09 cs.CL 79%

Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only Shallowly

Changjiang Gao, Hongda Hu, Peng Hu, Jiajun Chen, Jixing Li, Shujian Huang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01632 2024-04-03 cs.LG cs.SY eess.SY 79%

Enhancing Functional Safety in Automotive AMS Circuits through Unsupervised Machine Learning

Ayush Arunachalam, Ian Kintz, Suvadeep Banerjee, Arnab Raha, Xiankun Jin, Fei Su, Viswanathan Pillai Prasanth, Rubin A. Parekhji, Suriyaprakash Natarajan, Kanad Basu

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments 12 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18435 2024-03-28 cs.IR cs.CL 79%

DELTA: Pre-train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment

Haitao Li, Qingyao Ai, Xinyan Han, Jia Chen, Qian Dong, Yiqun Liu, Chong Chen, Qi Tian

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14149 2024-03-27 cs.CV cs.AI 79%

TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Qinying Liu, Wei Wu, Kecheng Zheng, Zhan Tong, Jiawei Liu, Yu Liu, Wei Chen, Zilei Wang, Yujun Shen

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏