arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8017 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8017 篇

1901.08872 2019-01-28 cs.NI 78%

Deep Learning-aided Application Scheduler for Vehicular Safety Communication

Mohammad Irfan Khan, François-Xavier Aubet, Marc-Oliver Pahl, Jérôme Härri

专题命中 其他安全 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.04344 2018-10-11 cs.RO 78%

Apprenticeship Bootstrapping Via Deep Learning with a Safety Net for UAV-UGV Interaction

Hung Nguyen, Vu Tran, Tung Nguyen, Matthew Garratt, Kathryn Kasmarik, Michael Barlow, Sreenatha Anavatti, Hussein Abbass

专题命中 其他安全 :safety(title,abstract)

Comments Presented at AI-HRI AAAI-FSS, 2018 (arXiv:1809.06606)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.04222 2018-05-22 cs.SI physics.soc-ph 78%

Graphlets versus node2vec and struc2vec in the task of network alignment

Shawn Gu, Tijana Milenkovic

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.00264 2018-02-02 cs.HC 78%

Automatic Safety Helmet Wearing Detection

Kang Li, Xiaoguang Zhao, Jiang Bian, Min Tan

专题命中 其他安全 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.02255 2017-10-09 cs.IR 78%

Fine-Grained Retrieval of Sports Plays using Tree-Based Alignment of Trajectories

Long Sha, Patrick Lucey, Stephan Zheng, Taehwan Kim, Yisong Yue, Sridha Sridharan

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.00546 2017-09-05 cs.RO 78%

Autonomous Waypoint Generation with Safety Guarantees: On-Line Motion Planning in Unknown Environments

Sanjeev Sharma

专题命中 其他安全 :safety(title,abstract)

Comments This paper was accepted for publication in the International Conference on Advanced Robotics 2013. It was not included in the final proceedings of the conference as I was unable to attend the conference to present the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.08062 2016-06-22 math.AP 78%

Global classical solutions of the Vlasov-Fokker-Planck equation with local alignment forces

Young-Pil Choi

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.00681 2016-02-17 physics.bio-ph q-bio.CB 78%

Mechanism for collective cell alignment in Myxococcus xanthus bacteria

Rajesh Balagam, Oleg A. Igoshin

专题命中 其他安全 :alignment(title,abstract)

Comments Added paragraph on high cell density simulations (new Supp. Figure S6) in Discussion section; Moved cell model and simulation procedure from Supplementary methods to Methods section in Main Text

详情

展开后加载摘要…

URL PDF HTML 收藏
1510.05301 2015-10-20 cs.SI cs.IR 78%

Social Media Analysis for Product Safety using Text Mining and Sentiment Analysis

Haruna Isah, Daniel Neagu, Paul Trundle

专题命中 其他安全 :safety(title,abstract)

Comments 2014 14th UK Workshop on Computational Intelligence (UKCI)

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.0991 2015-06-19 math.AP 78%

Critical thresholds in flocking hydrodynamics\\with nonlocal alignment

Eitan Tadmor, Changhui Tan

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1307.6569 2015-06-16 astro-ph.HE 78%

Alignment of supermassive black hole binary orbits and spins

M. Coleman Miller, Julian H. Krolik

专题命中 其他安全 :alignment(title,abstract)

Comments 18 pages, 1 figure. Accepted by The Astrophysical Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
1405.2417 2014-05-13 cs.NI 78%

Impact of Two Realistic Mobility Models for Vehicular Safety Applications

Md Habibur Rahman, Mohammad Nasiruddin

专题命中 其他安全 :safety(title,abstract)

Comments 6 pages, 9 figures, ICIEV 2014

Journal ref 3rd International Conference on Informatics, Electronics & Vision (ICIEV 2014)

详情

展开后加载摘要…

URL PDF HTML 收藏
1107.1623 2011-09-26 cond-mat.stat-mech q-bio.OT 78%

Mean-field theory of collective motion due to velocity alignment

Pawel Romanczuk, Lutz Schimansky-Geier

专题命中 其他安全 :alignment(title,abstract)

Comments corrected version, Ecological Complexity (2011) in press

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11221 2024-10-16 cs.LG cs.AI 77%

Multi-objective Reinforcement Learning: A Tool for Pluralistic Alignment

Peter Vamplew, Conor F Hayes, Cameron Foale, Richard Dazeley, Hadassah Harland

专题命中 其他安全 :alignment(title,comments);分类 cs.AI、cs.LG

Comments Accepted for the Pluralistic Alignment workshop at NeurIPS 2024. https://pluralistic-alignment.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06404 2025-06-10 cs.CL cs.AI cs.CY cs.LG 77%

Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights

Sooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao, Xing Xie, JinYeong Bak

机构 * Sungkyunkwan University(釜山大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01405 2025-03-04 cs.LG cs.AI cs.CL cs.CV cs.CY 77%

Representation Engineering: A Top-Down Approach to AI Transparency

Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, Dan Hendrycks

专题命中 其他安全 :safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Code is available at https://github.com/andyzoujm/representation-engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08809 2024-02-08 cs.CL 77%

Interpretability at Scale: Identifying Causal Mechanisms in Alpaca

Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, Noah D. Goodman

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL

Comments NeurIPS 2023 with Author Corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01478 2022-10-28 cs.CL cs.AI cs.CY cs.LG 77%

When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment

Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments NeurIPS 2022 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.02553 2018-01-03 cs.AI cs.SC 77%

Robust Computer Algebra, Theorem Proving, and Oracle AI

Gopal P. Sarma, Nick J. Hay

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments 15 pages, 3 figures

Journal ref Informatica Vol. 41 No. 3 (2017)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09937 2026-08-12 cs.CL cs.CY 新提交 76%

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

仔细考量文化:利用文化共识理论分析单文化与多文化场景下的大语言模型对齐

Krishna Pothugunta, John P. Lalor

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.CY

AI总结 本研究利用文化共识理论,分析大语言模型在单/多文化场景下的对齐情况,发现模型存在文化结构误表征问题,该理论可用于区分模型反映人类多样性与算法同质化的情况。

Comments Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17629 2026-07-28 cs.LG cs.AI 76%

Learning Chemical Reaction Representation with Reactant-Product Alignment

基于反应物-产物对齐的学习化学反应表示

Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出RAlign模型,通过反应物-产物对齐和反应中心感知注意力机制,提升化学反应表示学习的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09936 2026-07-14 cs.LG cs.AI cs.CR 新提交 76%

SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification

SMETA-ZSL:用于零样本威胁分类的语义元对齐

Ivan Alejandro Montoya Sanchez, Anantaa Kotal, Aritran Piplai

机构 * The University of Texas at El Paso(德克萨斯大学艾尔帕索分校)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 研究针对网络安全新威胁无标注数据问题,提出SMETA-ZSL方法,通过对比微调、情景元学习和知识蒸馏等,从重叠语言描述学习语义原型并对齐行为特征,实现跨可见-未见类别的泛化,在7个基准测试中性能远超先前方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01592 2026-07-14 cs.CY cs.CL 交叉投稿 76%

Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises

问题类型、认知负荷与CEFR对齐:评估LLM生成的EFL语法练习

Steve Woollaston, Brendan Flanagan, Yuko Toyokawa, Hiroaki Ogata

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.CY

AI总结 本研究通过分析日本初中生在语法练习应用中的日志数据,评估了LLM生成的EFL学习内容的教学可行性,揭示了不同问题模态对表现的影响,并验证了CEFR-J语法框架的难度层级。

Comments Under review for the the 34th International Conference on Computers in Education (ICCE 2026). 2jun26: v2 - fixed minor typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31394 2026-07-03 cs.LG cs.AI cs.CV q-bio.QM 新提交 76%

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

解决AI中的叠加问题以实现可解释性与患者-神经元图像的跨模态对齐

Jisung Park, Seohyeon Kang, Daeun Yoo, Eunsu Lee, Seoin Cho, Wooyeop Choi, Ian Choi, James R. Evan, Daesoo Kim, Sonia Gandhi, Minee L. Choi

机构 * KAIST(韩国科学技术院) Konyang University(建阳大学) Chang Gung University(长庚大学) UCL Queen Square Institute of Neurology & The Francis Crick Institute(伦敦大学学院皇后广场神经病学研究所与弗朗西斯·克里克研究所)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 利用稀疏自编码器解决高维生物数据中神经网络表示空间的叠加问题,恢复几何保真度,并通过Gromov-Wasserstein最优传输实现图像与单细胞RNA测序数据的跨模态对齐。

Comments 10 pages, 7 figures (plus 14 in appendix), 1 table, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04391 2026-07-03 cs.AI cs.CL cs.SI q-bio.NC 版本更新 76%

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

心理想象网络显示人类跨群体中心性和聚类对齐,而大型语言模型无法复制

Saurabh Ranjan, Brian Odegaard

机构 * University of Florida(佛罗里达大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本研究通过心理网络分析发现,人类对心理意象的生动性评分在不同文化群体中形成稳定的网络结构,而大型语言模型(LLM)无法复制这种结构,表明人类想象网络根植于具身经验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08016 2026-06-09 cs.CV cs.AI cs.CL 新提交 76%

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

IEA:通过三阶段多任务对齐的业余友好型对话式图像编辑代理

Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institution(上海创新研究院) Huawei Technologies Ltd.(华为技术有限公司) Nanyang Technological University(南洋理工大学) Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 提出IEA对话式图像编辑代理,通过三阶段多任务训练学习操作参数化工具,实现可解释编辑轨迹,在像素距离和ROUGE-L指标上优于基线,用户研究中指令跟随和感知质量表现最佳。

Comments [CVPR 2026 Findings] Our data and code are released at https://github.com/OpenDFM/Image_Edit_Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29458 2026-05-29 cs.CL cs.AI 76%

Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

面向LLM人格模拟的自适应访谈:基于证据的推理提升决策对齐

Ruoxi Su, Yuhan Liu, Jingyu Hu

机构 * University of Cambridge(剑桥大学) Independent Researcher(独立研究员)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 提出自适应访谈框架,通过结构化三阶段对话收集人格相关信息,并基于访谈记录评估LLM在道德困境场景中模拟个体决策的能力,发现基于后续追问的证据推理能显著提升预测准确性。

Comments 20 pages, 2 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25189 2026-05-26 cs.LG cs.CL 76%

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

方向对齐缓解语言模型强化学习中的奖励黑客问题

Wenlong Deng, Jiaji Huang, Kaan Ozkara, Yushu Li, Christos Thrampoulidis, Xiaoxiao Li, Youngsuk Park

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Amazon(亚马逊)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

AI总结 通过分析强化学习更新的几何结构,发现奖励黑客源于优化偏离稳定低维学习轨迹,提出可信方向投影方法约束梯度在干净参考子空间内,延迟捷径利用并保持任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20730 2026-05-21 cs.CL cs.AI 76%

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning

分布对齐作为设计任务向量在上下文学习中的准则

Jihoon Kwon, Jiwon Choi, Jy-yong Sohn

机构 * Seoul National University(首尔国立大学) Yonsei University(延世大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出通过分布对齐来设计任务向量,引入了NTP距离作为衡量指标,并开发了线性任务向量方法以提升性能和效率。

Comments 9 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12910 2026-04-23 cs.CL cs.AI 76%

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment

SciCoQA:科学论文与代码对齐的质量保障

Tim Baumgärtner, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普遍知识处理实验室(UKP实验室)、计算机科学系、德累斯顿技术大学和应用网络安全国家研究中心ATHENE)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出SciCoQA数据集,用于评估LLM在科学论文与代码对齐任务中的表现,揭示自动化科学质量保障的关键差距。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏