arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9419 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9419 篇

2406.13663 2024-12-02 cs.CL cs.AI cs.LG 78%

Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation

Jirui Qi, Gabriele Sarti, Raquel Fernández, Arianna Bisazza

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by EMNLP 2024 Main Conference. Code and data released at https://github.com/Betswish/MIRAGE

Journal ref Proceedings of EMNLP (2024) 6037-6053

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17981 2024-11-28 cs.SE 78%

Engineering Trustworthy Software: A Mission for LLMs

Marco Vieira

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15735 2024-11-26 cs.CV 78%

Test-time Alignment-Enhanced Adapter for Vision-Language Models

Baoshun Tong, Kaiyu Song, Hanjiang Lai

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06099 2024-11-12 cs.HC 78%

CoPrompter: User-Centric Evaluation of LLM Instruction Alignment for Improved Prompt Engineering

Ishika Joshi, Simra Shahid, Shreeya Venneti, Manushree Vasu, Yantao Zheng, Yunyao Li, Balaji Krishnamurthy, Gromit Yeuk-Yin Chan

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23207 2024-11-04 eess.SY cs.SY 78%

Enhancing Autonomous Driving Safety Analysis with Generative AI: A Comparative Study on Automated Hazard and Risk Assessment

Alireza Abbaspour, Aliasghar Arab, Yashar Mousavi

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23988 2024-11-01 cs.CV 78%

JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment

Joao Sousa, Roya Darabi, Armando Sousa, Frank Brueckner, Luís Paulo Reis, Ana Reis

专题命中 安全评测 :alignment(title,abstract)

Comments 26 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21736 2024-10-31 cs.RO 78%

Enhancing Safety and Robustness of Vision-Based Controllers via Reachability Analysis

Kaustav Chakraborty, Aryaman Gupta, Somil Bansal

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18927 2024-10-25 cs.CR 78%

SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models

Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, Dacheng Tao

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13477 2024-10-18 cs.DC cs.CR 78%

Advocate -- Trustworthy Evidence in Cloud Systems

Sebastian Werner, Sepideh Masoudi, Fernando Castillo, Fabian Piper, Jonathan Heiss

专题命中 安全评测 :trustworthy(title,abstract)

Comments Preprint version of the paper at 6th Conference on Blockchain Research & Applications for Innovative Networks and Services (BRAINS'24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12475 2024-10-18 cs.MA 78%

Aegis:An Advanced LLM-Based Multi-Agent for Intelligent Functional Safety Engineering

Lu Shi, Bin Qi, Jiarui Luo, Yang Zhang, Zhanzhao Liang, Zhaowei Gao, Wenke Deng, Lin Sun

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14122 2024-10-18 cs.CV cs.CR 78%

SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution

Zhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng, Qinglong Wang, Zhan Qin, Zhibo Wang, Kui Ren

专题命中 安全评测 :safety(title,abstract)

Comments To appear in the the 31st ACM Conference on Computer and Communications Security (CCS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12225 2024-10-17 cs.CV 78%

Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety

Lucas Choi, Ross Greer

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18878 2024-10-07 cs.CL cs.AI cs.CY cs.IR 78%

Suicide Phenotyping from Clinical Notes in Safety-Net Psychiatric Hospital Using Multi-Label Classification with Pre-Trained Language Models

Zehan Li, Yan Hu, Scott Lane, Salih Selek, Lokesh Shahani, Rodrigo Machado-Vieira, Jair Soares, Hua Xu, Hongfang Liu, Ming Huang

专题命中 安全评测 :safety(title);分类 cs.CL、cs.AI、cs.CY

Comments submitted to AMIA Informatics Summit 2025 as a conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16033 2024-09-25 cs.RO 78%

RTAGrasp: Learning Task-Oriented Grasping from Human Videos via Retrieval, Transfer, and Alignment

Wenlong Dong, Dehao Huang, Jiangshan Liu, Chao Tang, Hong Zhang

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14833 2024-09-24 cs.RO 78%

SymAware: A Software Development Framework for Trustworthy Multi-Agent Systems with Situational Awareness

Ernesto Casablanca, Zengjie Zhang, Gregorio Marchesini, Sofie Haesaert, Dimos V. Dimarogonas, Sadegh Soudjani

专题命中 安全评测 :trustworthy(title,abstract)

Comments Submitted to ICRA2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21654 2024-08-01 cs.CV 78%

MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

Anurag Das, Xinting Hu, Li Jiang, Bernt Schiele

专题命中 安全评测 :alignment(title,abstract)

Comments accepted at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19708 2024-07-26 cs.LG cs.AI cs.CL cs.HC 78%

Harmonic LLMs are Trustworthy

Nicholas S. Kersting, Mohammad Rahman, Suchismitha Vedala, Yang Wang

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.LG

Comments 15 pages, 2 figures, 16 tables; added Claude-3.0, GPT-4o, Mistral-7B, Mixtral-8x7B, and more annotation for other models

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11387 2024-07-17 cs.HC 78%

A Framework for Evaluating Appropriateness, Trustworthiness, and Safety in Mental Wellness AI Chatbots

Lucia Chen, David A. Preece, Pilleriin Sikka, James J. Gross, Ben Krause

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06555 2024-07-10 cs.CL cs.AI cs.CV cs.LG 78%

Do Vision and Language Models Share Concepts? A Vector Space Alignment Study

Jiaang Li, Yova Kementchedjhieva, Constanza Fierro, Anders Søgaard

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI、cs.LG

Comments 12 pages, long paper accepted by TACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00979 2024-07-02 cs.CV 78%

Cross-Modal Attention Alignment Network with Auxiliary Text Description for zero-shot sketch-based image retrieval

Hanwen Su, Ge Song, Kai Huang, Jiyan Wang, Ming Yang

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17600 2024-06-21 cs.CV 78%

MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, Yu Qiao

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19009 2024-06-17 cs.CV 78%

Enhancing Vision-Language Model with Unmasked Token Alignment

Jihao Liu, Jinliang Zheng, Boxiao Liu, Yu Liu, Hongsheng Li

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted by TMLR; Code and models are available at https://github.com/jihaonew/UTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12933 2024-06-04 cs.CL cs.AI cs.LG 78%

Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs

Bilgehan Sel, Priya Shanmugasundaram, Mohammad Kachuee, Kun Zhou, Ruoxi Jia, Ming Jin

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2024, long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14113 2024-05-24 eess.IV cs.CV 78%

Multi-modality Regional Alignment Network for Covid X-Ray Survival Prediction and Report Generation

Zhusi Zhong, Jie Li, John Sollee, Scott Collins, Harrison Bai, Paul Zhang, Terrence Healey, Michael Atalay, Xinbo Gao, Zhicheng Jiao

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05340 2024-05-10 cs.SE 78%

POLARIS: A framework to guide the development of Trustworthy AI systems

Maria Teresa Baldassarre, Domenico Gigante, Marcos Kalinowski, Azzurra Ragone

专题命中 安全评测 :trustworthy(title,abstract)

Journal ref Conference on AI Engineering Software Engineering for AI (CAIN 2024), April 14--15, 2024, Lisbon, Portugal

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12814 2024-04-23 cs.HC 78%

Exploring the Impact of AI Value Alignment in Collaborative Ideation: Effects on Perception, Ownership, and Output

Alicia Guo, Pat Pataranutaporn, Pattie Maes

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11249 2024-04-18 cs.CV 78%

A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene

Wenbo Zhang, Yifan Zhang, Jianfeng Lin, Binqiang Huang, Jinlu Zhang, Wenhao Yu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09555 2024-04-16 cs.CV 78%

AI-KD: Towards Alignment Invariant Face Image Quality Assessment Using Knowledge Distillation

Žiga Babnik, Fadi Boutros, Naser Damer, Peter Peer, Vitomir Štruc

专题命中 安全评测 :alignment(title,abstract)

Comments IEEE International Workshop on Biometrics and Forensics (IWBF) 2024, pp. 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07479 2024-04-12 cs.HC 78%

RASSAR: Room Accessibility and Safety Scanning in Augmented Reality

Xia Su, Han Zhang, Kaiming Cheng, Jaewook Lee, Qiaochu Liu, Wyatt Olson, Jon Froehlich

专题命中 安全评测 :safety(title,abstract)

Comments To Appear in CHI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06471 2024-03-12 cs.CV 78%

Toward Robust Canine Cardiac Diagnosis: Deep Prototype Alignment Network-Based Few-Shot Segmentation in Veterinary Medicine

Jun-Young Oh, In-Gyu Lee, Tae-Eui Kam, Ji-Hoon Jeong

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏