arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

1905.01320 2019-05-07 cs.LG cs.AI stat.ML 62%

Meta-learners' learning dynamics are unlike learners'

Neil C. Rabinowitz

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 26 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.05659 2019-02-21 cs.CV cs.AI cs.HC cs.LG 62%

Artificial Intelligence Assisted Infrastructure Assessment Using Mixed Reality Systems

Enes Karaaslan, Ulas Bagci, F. Necati Catbas

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 5,240 word texts, 3 tables, 14 figures. Transportation Research Record: Journal of the Transportation Research Board, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.01813 2018-12-06 cs.CY cs.LG 62%

Machine-learned epidemiology: real-time detection of foodborne illness at scale

Adam Sadilek, Stephanie Caty, Lauren DiPrete, Raed Mansour, Tom Schenk, Mark Bergtholdt, Ashish Jha, Prem Ramaswami, Evgeniy Gabrilovich

专题命中 其他安全 :safety(abstract);分类 cs.CY、cs.LG

Journal ref npj Digital Medicine 1:36 (2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.10284 2018-11-28 cs.AI cs.LG 62%

Enter the Matrix: Safely Interruptible Autonomous Systems via Virtualization

Mark O. Riedl, Brent Harrison

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages; 1 figure; title, abstract updated; new experimental results

Journal ref Proceedings of the AAAI 2019 Workshop on SafeAI

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.04587 2018-11-21 cs.LG cs.AI cs.NE stat.ML 62%

Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures

Sergey Bartunov, Adam Santoro, Blake A. Richards, Luke Marris, Geoffrey E. Hinton, Timothy Lillicrap

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments NIPS 2018. Version 2 contains more experimental data including best hyperparameters found

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.05544 2018-11-15 cs.CL cs.LG stat.ML 62%

An Introductory Survey on Attention Mechanisms in NLP Problems

Dichao Hu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.02724 2018-10-08 cs.GL cs.AI cs.CY 62%

Human Indignity: From Legal AI Personhood to Selfish Memes

Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.01989 2018-10-05 cs.AI cs.LG 62%

Verification for Machine Learning, Autonomy, and Neural Networks Survey

Weiming Xiang, Patrick Musau, Ayana A. Wild, Diego Manzanas Lopez, Nathaniel Hamilton, Xiaodong Yang, Joel Rosenfeld, Taylor T. Johnson

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07899 2018-08-27 cs.AI cs.CY 62%

A Century Long Commitment to Assessing Artificial Intelligence and its Impact on Society

Barbara J. Grosz, Peter Stone

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

Comments This paper will appear in Communications of the ACM (December 2018, vol 61, no. 12)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.09938 2018-05-28 cs.AI cs.LG 62%

Automated Verification of Neural Networks: Advances, Challenges and Perspectives

Francesco Leofante, Nina Narodytska, Luca Pulina, Armando Tacchella

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00900 2018-05-03 cs.AI cs.CL cs.CV cs.IR 62%

Images & Recipes: Retrieval in the cooking context

Micael Carvalho, Rémi Cadène, David Picard, Laure Soulier, Matthieu Cord

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Published at DECOR / ICDE 2018. Extended version accepted at SIGIR 2018, available here: arXiv:1804.11146

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.01096 2018-02-09 cs.SE cs.AI cs.HC cs.LG cs.NE 62%

Software Engineers vs. Machine Learning Algorithms: An Empirical Study Assessing Performance and Reuse Tasks

Nathalia Nascimento, Carlos Lucena, Paulo Alencar, Donald Cowan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 22 pages. To be submitted to IEEE Transactions on Software Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.01996 2017-12-07 eess.AS cs.AI cs.CL cs.SD 62%

An analysis of incorporating an external language model into a sequence-to-sequence model

Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N. Sainath, Zhifeng Chen, Rohit Prabhavalkar

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.00177 2017-09-25 cs.CL cs.AI cs.NE 62%

LSTM Neural Reordering Feature for Statistical Machine Translation

Yiming Cui, Shijin Wang, Jianfeng Li

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6 pages, accepted by NAACL2016 short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.08535 2017-07-17 cs.SE cs.CL cs.LG cs.NE 62%

Deep API Learning

Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, Sunghun Kim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments The paper is accepted at FSE 2016 (the 24th ACM SIGSOFT International Symposium on the Foundations of Software Engineering)

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.01086 2017-03-28 cs.AI cs.LG cs.RO 62%

Deep Learning of Robotic Tasks without a Simulator using Strong and Weak Human Supervision

Bar Hilleli, Ran El-Yaniv

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.3427 2014-01-16 cs.LG cs.AI 62%

Analogical Dissimilarity: Definition, Algorithms and Two Experiments in Machine Learning

Laurent Miclet, Sabri Bayoudh, Arnaud Delhay

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref Journal Of Artificial Intelligence Research, Volume 32, pages 793-824, 2008

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13254 2026-06-12 cs.CL 新提交 61%

Evaluating Pluralism in LLMs through Latent Perspectives

通过潜在视角评估LLM中的多元主义

Laura Majer, Jan Šnajder, Martin Tutek

机构 * University of Helsinki(赫尔辛基大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL

AI总结 提出一种领域无关的多层无监督框架,从LLM生成文本中提取潜在视角,评估多元主义差距,发现稀有视角仍被不成比例地低估。

Comments Pluralistic Alignment Workshop @ ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10515 2026-03-23 physics.comp-ph cs.CE cs.LG cs.SY eess.SY 61%

Virtual Sensing for Solder Layer Degradation and Temperature Monitoring in IGBT Modules

IGBT模块焊层退化与温度监测的虚拟传感

Andrea Urgolo, Monika Stipsitz, Hèlios Sanchis-Alepuz

机构 * Silicon Austria Labs GmbH(硅 Austria 实验室)

专题命中 其他安全 :safety(abstract,journal_ref);分类 cs.LG

AI总结 本文利用机器学习虚拟传感技术,通过有限物理传感器估计IGBT模块焊层退化状态及温度分布,实现高精度退化区域估计和表面温度复现。

Comments Andrea Urgolo and Monika Stipsitz contributed equally to this work

Journal ref 2025 9th International Conference on System Reliability and Safety (ICSRS), Turin, Italy, 2025, pp. 538-547

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09020 2026-03-11 cs.HC cs.AI 61%

AI Phenomenology for Understanding Human-AI Experiences Across Eras

人工智能现象学:理解跨时代的以人为本的AI体验

Bhada Yun, Evgenia Taranova, Dana Feng, Renn Su, April Yi Wang

机构 * ETH Zürich(苏黎世联邦理工学院) University of Bergen(卑尔根大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI

AI总结 本文提出人工智能现象学,通过研究用户与AI交互的主观体验,促进双向人机对齐,并提供可重复的方法论工具。

Comments This is an accepted workshop paper at CHI '26, "W37: Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures", or https://bialign-workshop.github.io/2026/cfp

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03074 2026-03-04 cs.HC cs.AI 61%

Design Generative AI for Practitioners: Exploring Interaction Approaches Aligned with Creative Practice

为实践者设计生成式AI:探索与创意实践对齐的交互方法

Xiaohan Peng, Wendy E. Mackay, Janin Koch

机构 * LISN Université Paris-Saclay, CNRS, Inria(LISN 巴黎-萨克雷大学,法国国家科学研究中心,法国国家信息与自动化技术研究所) UMR 9189 CRIStAL Univ. Lille, Inria, CNRS, Centrale Lille(UMR 9189 CRIStAL 利摩日大学,法国国家信息与自动化技术研究所,法国国家科学研究中心,利摩日中央理工大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI

AI总结 本文提出三种交互方法,帮助设计师在不同阶段引导生成式AI与创意实践对齐,强调动态协商与主动/被动角色的适应性。

Comments Accepted to ACM CHI 2026 Workshop on Bidirectional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18745 2025-10-22 cs.CL 61%

Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting

Taha Binhuraib, Greta Tuckute, Nicholas Blauch

机构 * Novus Technologies MIT(麻省理工学院) Harvard University(哈佛大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL

Comments ICLR 2024 Workshop on Representational Alignment (Re-Align) Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14390 2025-06-18 cs.LG cs.CV 61%

Enclosing Prototypical Variational Autoencoder for Explainable Out-of-Distribution Detection

Conrad Orglmeister, Erik Bochinski, Volker Eiselein, Elvira Fleig

机构 * Digitale Schiene Deutschland, DB InfraGO AG(Digitale Schiene Deutschland,DB InfraGO AG) Communication Systems Group, Technische Universität Berlin(通信系统组,技术大学柏林)

专题命中 其他安全 :safety(abstract,comments);分类 cs.LG

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Computer Safety, Reliability and Security - SAFECOMP 2024 Workshops - DECSoS, SASSUR, TOASTS, and WAISE, and is available online at https://doi.org/10.1007/978-3-031-68738-9_29

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11096 2025-03-17 cs.CV cs.AI cs.HC 61%

Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation

He Zhang, Xinyi Fu, John M. Carroll

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI

Comments This paper will appear at ICLR 2025 Workshop on Bidirectional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14095 2025-02-21 cs.CL 61%

Retrieving Versus Understanding Extractive Evidence in Few-Shot Learning

Karl Elbakian, Samuel Carton

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL

Comments 9 pages, 8 figures, Accepted to AAAI 2025 Main Conference (AI Alignment Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06204 2024-03-12 cs.CL 61%

Identifying and interpreting non-aligned human conceptual representations using language modeling

Wanqian Bao, Uri Hasson

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL

Comments To appear at the ICLR 2024 Workshop on Representational Alignment (Re-Align)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.08138 2023-01-20 cs.SE cs.AI cs.SY eess.SY 61%

Architecting Safer Autonomous Aviation Systems

Jane Fenn, Mark Nicholson, Ganesh Pai, Michael Wilkinson

专题命中 其他安全 :safety(abstract,comments);分类 cs.AI

Comments 18 pages, 12 figures, to appear in the proceedings of the 2023 Safety-critical Systems Symposium (SSS '23), York, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09959 2022-10-19 cs.LG 61%

Out of Distribution Reasoning by Weakly-Supervised Disentangled Logic Variational Autoencoder

Zahra Rahiminasab, Michael Yuhas, Arvind Easwaran

专题命中 其他安全 :safety(abstract,comments);分类 cs.LG

Comments Accepted in The 6th International Conference on System Reliability and Safety (ICSRS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12991 2022-03-01 cs.AI cs.CR 61%

Attacks and Faults Injection in Self-Driving Agents on the Carla Simulator -- Experience Report

Niccolò Piazzesi, Massimo Hong, Andrea Ceccarelli

专题命中 其他安全 :safety(abstract,comments);分类 cs.AI

Comments submitted version; appeared at: International Conference on Computer Safety, Reliability, and Security. Springer, Cham, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00431 2020-03-03 cs.AI 61%

A Study on Multimodal and Interactive Explanations for Visual Question Answering

Kamran Alipour, Jurgen P. Schulze, Yi Yao, Avi Ziskind, Giedrius Burachas

专题命中 其他安全 :safety(abstract,journal_ref);分类 cs.AI

Comments http://ceur-ws.org/Vol-2560/paper44.pdf

Journal ref Proceedings of the Workshop on Artificial Intelligence Safety (SafeAI 2020) co-located with 34th AAAI Conference on Artificial Intelligence (AAAI 2020), New York, USA, Feb 7, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏