arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2406.10297 2024-06-18 cs.CL cs.AI 62%

SememeLM: A Sememe Knowledge Enhanced Method for Long-tail Relation Representation

Shuyi Li, Shaojuan Wu, Xiaowang Zhang, Zhiyong Feng

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17053 2024-06-18 cs.NI cs.AI cs.LG 62%

WirelessLLM: Empowering Large Language Models Towards Wireless Intelligence

Jiawei Shao, Jingwen Tong, Qiong Wu, Wei Guo, Zijian Li, Zehong Lin, Jun Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06624 2024-06-17 cs.LG cs.CY 62%

Exploring the Determinants of Pedestrian Crash Severity Using an AutoML Approach

Amir Rafe, Patrick A. Singleton

专题命中 其他安全 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14492 2024-06-17 cs.CL cs.AI 62%

Towards Robust Instruction Tuning on Multimodal Large Language Models

Wei Han, Hui Chen, Soujanya Poria

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04268 2024-06-07 cs.LG cs.AI 62%

Open-Endedness is Essential for Artificial Superhuman Intelligence

Edward Hughes, Michael Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, Tim Rocktaschel

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00160 2024-06-07 cs.CL cs.AI 62%

Self-Specialization: Uncovering Latent Expertise within Large Language Models

Junmo Kang, Hongyin Luo, Yada Zhu, Jacob Hansen, James Glass, David Cox, Alan Ritter, Rogerio Feris, Leonid Karlinsky

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2024 (Findings; Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03030 2024-06-06 cs.CL cs.LG 62%

From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation

Ali Malik, Stephen Mayhew, Chris Piech, Klinton Bicknell

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Journal ref In Findings of the Association for Computational Linguistics (ACL 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02600 2024-06-06 cs.LG cs.AI stat.ML 62%

Data Quality in Edge Machine Learning: A State-of-the-Art Survey

Mohammed Djameleddine Belgoumri, Mohamed Reda Bouadjenek, Sunil Aryal, Hakim Hacid

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02030 2024-06-06 cs.CL cs.AI 62%

Multimodal Reasoning with Multimodal Knowledge Graph

Junlin Lee, Yequan Wang, Jing Li, Min Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18721 2024-06-06 cs.CV cs.AI cs.CL 62%

Correctable Landmark Discovery via Large Models for Vision-Language Navigation

Bingqian Lin, Yunshuang Nie, Ziming Wei, Yi Zhu, Hang Xu, Shikui Ma, Jianzhuang Liu, Xiaodan Liang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by TPAMI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01382 2024-06-04 cs.CL cs.AI 62%

Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function

Keyon Vafa, Ashesh Rambachan, Sendhil Mullainathan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments To appear in ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.21068 2024-06-03 cs.CL cs.AI 62%

Code Pretraining Improves Entity Tracking Abilities of Language Models

Najoung Kim, Sebastian Schuster, Shubham Toshniwal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19550 2024-05-31 cs.LG cs.CL 62%

Stress-Testing Capability Elicitation With Password-Locked Models

Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, David Krueger

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03893 2024-05-31 cs.CL cs.AI 62%

From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models

Luiza Pozzobon, Patrick Lewis, Sara Hooker, Beyza Ermis

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15789 2024-05-28 cs.AI cs.IT cs.LG cs.LO math.IT 62%

Semantic Objective Functions: A distribution-aware method for adding logical constraints in deep learning

Miguel Angel Mendez-Lucero, Enrique Bojorquez Gallardo, Vaishak Belle

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09647 2024-05-24 cs.SI cs.CL cs.CY 62%

Large Language Models Help Reveal Unhealthy Diet and Body Concerns in Online Eating Disorders Communities

Minh Duc Chu, Zihao He, Rebecca Dorn, Kristina Lerman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13219 2024-05-24 cs.AI cs.CL 62%

How Reliable AI Chatbots are for Disease Prediction from Patient Complaints?

Ayesha Siddika Nipu, K M Sajjadul Islam, Praveen Madiraju

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 24th IEEE International Conference on Information Reuse and Integration (IEEE IRI 2024), San Jose, CA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02752 2024-05-22 cs.LG cs.AI 62%

Offline Reinforcement Learning with Imbalanced Datasets

Li Jiang, Sijie Cheng, Jielin Qiu, Haoran Xu, Wai Kin Chan, Zhao Ding

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref ICML 2023, workshop on Data-centric Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14381 2024-05-14 cs.CL cs.AI 62%

Editing Knowledge Representation of Language Model via Rephrased Prefix Prompts

Yuchen Cai, Ding Cao, Rongxi Guo, Yaqin Wen, Guiquan Liu, Enhong Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19pages,3figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14236 2024-05-14 cs.RO cs.AI cs.CV cs.LG 62%

MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation

Patrick Lancaster, Nicklas Hansen, Aravind Rajeswaran, Vikash Kumar

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02795 2024-05-07 cs.AI cs.CL 62%

Evaluating and Optimizing Educational Content with Large Language Model Judgments

Joy He-Yueya, Noah D. Goodman, Emma Brunskill

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12420 2024-05-07 cs.CL cs.AI 62%

LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models

Shizhe Diao, Rui Pan, Hanze Dong, Ka Shun Shum, Jipeng Zhang, Wei Xiong, Tong Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Published in NAACL 2024 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02105 2024-05-06 cs.AI cs.CL cs.IT math.IT 62%

Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph

Vladyslav Nechakhin, Jennifer D'Souza, Steffen Eger

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 11 figures. In review at https://www.mdpi.com/journal/information/special_issues/WYS02U2GTD

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01561 2024-05-06 cs.SE cs.AI cs.CY 62%

Rapid Mobile App Development for Generative AI Agents on MIT App Inventor

Jaida Gao, Calab Su, Etai Miller, Kevin Lu, Yu Meng

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

Journal ref Journal of advances in information science and technology 2(3) 1-8, March 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15597 2024-04-25 cs.NE cs.AI cs.LG cs.MA 62%

GRSN: Gated Recurrent Spiking Neurons for POMDPs and MARL

Lang Qin, Ziming Wang, Runhao Jiang, Rui Yan, Huajin Tang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12399 2024-04-25 cs.LG cs.CL cs.SI 62%

A Survey of Graph Meets Large Language Model: Progress and Future Directions

Yuhan Li, Zhixun Li, Peisong Wang, Jia Li, Xiangguo Sun, Hong Cheng, Jeffrey Xu Yu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments IJCAI 2024 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14906 2024-04-24 cs.CV cs.AI cs.LG 62%

Driver Activity Classification Using Generalizable Representations from Vision-Language Models

Ross Greer, Mathias Viborg Andersen, Andreas Møgelmose, Mohan Trivedi

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14864 2024-04-23 cs.AI cs.CV cs.LG 62%

Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability

Georgii Mikriukov, Gesina Schwalbe, Christian Hellert, Korinna Bade

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11797 2024-04-19 cs.CV cs.AI cs.LG 62%

When are Foundation Models Effective? Understanding the Suitability for Pixel-Level Classification Using Multispectral Imagery

Yiqun Xie, Zhihao Wang, Weiye Chen, Zhili Li, Xiaowei Jia, Yanhua Li, Ruichen Wang, Kangyang Chai, Ruohan Li, Sergii Skakun

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11589 2024-04-18 cs.CV cs.AI cs.LG 62%

Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding

Zezhong Fan, Xiaohan Li, Chenhao Fang, Topojoy Biswas, Kaushiki Nag, Jianpeng Xu, Kannan Achan

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments WWW 2024 Companion

详情

展开后加载摘要…

URL PDF HTML 收藏