arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8033 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8033 篇

2305.11952 2023-05-23 cs.CL 74%

Self-QA: Unsupervised Knowledge Guided Language Model Alignment

Xuanyu Zhang, Qing Yang

专题命中 其他安全 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08535 2022-11-10 cs.CL 74%

The challenges of temporal alignment on Twitter during crises

Aniket Pramanick, Tilman Beck, Kevin Stowe, Iryna Gurevych

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Accepted to Findings of EMNLP, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04482 2022-03-31 cs.CV cs.CL 74%

FLAVA: A Foundational Language And Vision Alignment Model

Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, Douwe Kiela

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14133 2022-01-13 cs.LG 74%

Self-Labeling of Fully Mediating Representations by Graph Alignment

Martijn Oldenhof, Adam Arany, Yves Moreau, Jaak Simm

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments Code available: https://github.com/biolearning-stadius/chemgrapher-self-rich-labeling

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.05862 2021-11-22 cs.LG stat.ML 74%

Constrained Non-Affine Alignment of Embeddings

Yuwei Wang, Yan Zheng, Yanqing Peng, Chin-Chia Michael Yeh, Zhongfang Zhuang, Das Mahashweta, Bendre Mangesh, Feifei Li, Wei Zhang, Jeff M. Phillips

专题命中 其他安全 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04316 2021-11-10 cs.CV cs.LG cs.NE 74%

COVID-19 Face Mask Recognition with Advanced Face Cut Algorithm for Human Safety Measures

Arkaprabha Basu, Md Firoj Ali

专题命中 其他安全 :safety(title);分类 cs.LG

Comments 5 pages, 7 figures

Journal ref 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12028 2021-09-27 cs.CL 74%

Investigating Post-pretraining Representation Alignment for Cross-Lingual Question Answering

Fahim Faisal, Antonios Anastasopoulos

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Accepted at MRQA Workshop 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12547 2021-04-13 cs.CL 74%

Multilingual BERT Post-Pretraining Alignment

Lin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, Mo Yu

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Accepted at NAACL2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00195 2020-12-02 cs.LG q-bio.BM 74%

Profile Prediction: An Alignment-Based Pre-Training Task for Protein Sequence Models

Pascal Sturmfels, Jesse Vig, Ali Madani, Nazneen Fatema Rajani

专题命中 其他安全 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.03590 2020-11-30 cs.RO cs.LG cs.SY eess.SY 74%

Reactive motion planning with probabilistic safety guarantees

Yuxiao Chen, Ugo Rosolia, Chuchu Fan, Aaron D. Ames, Richard Murray

专题命中 其他安全 :safety(title);分类 cs.LG

Comments In the Conference on Robotic Learning 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.01056 2020-01-07 cs.LG stat.AP stat.ML 74%

Root Cause Detection Among Anomalous Time Series Using Temporal State Alignment

Sayan Chakraborty, Smit Shah, Kiumars Soltani, Anna Swigart

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments 6 pages, 7 figures, 2019 18th IEEE International Conference on Machine Learning and Applications (ICMLA)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.05727 2019-11-14 cs.CY cs.IR eess.IV 74%

Artificial Intelligence Strategies for National Security and Safety Standards

Erik Blasch, James Sung, Tao Nguyen, Chandra P. Daniel, Alisa P. Mason

专题命中 其他安全 :safety(title);分类 cs.CY

Comments Presented at AAAI FSS-19: Artificial Intelligence in Government and Public Sector, Arlington, Virginia, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.03562 2019-05-22 cs.LG stat.ML 74%

Real time Traffic Flow Parameters Prediction with Basic Safety Messages at Low Penetration of Connected Vehicles

Mizanur Rahman, Mashrur Chowdhury, Jerome McClendon

专题命中 其他安全 :safety(title);分类 cs.LG

Comments 16 pages, 15 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.08129 2018-11-21 cs.IR cs.CL 74%

Alignment Analysis of Sequential Segmentation of Lexicons to Improve Automatic Cognate Detection

Pranav A

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Published at ACL-SRW 2018

Journal ref Proceedings of ACL 2018, Student Research Workshop. 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.03479 2018-03-12 cs.AI 74%

Highly Automated Learning for Improved Active Safety of Vulnerable Road Users

Maarten Bieshaar, Günther Reitberger, Viktor Kreß, Stefan Zernetsch, Konrad Doll, Erich Fuchs, Bernhard Sick

专题命中 其他安全 :safety(title);分类 cs.AI

Comments 4 pages, 1 figure

Journal ref published in ACM Chapters Computer Science in Cars Symposium (CSCS-17). Munich, Germany. 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.07398 2016-08-29 cs.LO cs.AI cs.SE 74%

Proceedings First Workshop on Causal Reasoning for Embedded and safety-critical Systems Technologies

Gregor Gössler, Oleg Sokolsky

专题命中 其他安全 :safety(title);分类 cs.AI

Journal ref EPTCS 224, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.8053 2014-09-30 cs.AI 74%

Medical diagnosis as pattern recognition in a framework of information compression by multiple alignment, unification and search

J. Gerard Wolff

专题命中 其他安全 :alignment(title);分类 cs.AI

Journal ref Decision Support Systems 42, 608-625, 2006

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23623 2026-08-26 cs.SE cs.AI cs.LG 新提交 73%

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

智能体何时可以停止?带证据的工具使用大语言模型终止机制

Jason Liu

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 该研究针对工具使用大语言模型的终止问题,提出带证据的终止(ECT)机制,经实验验证其能显著减少不安全完成与过早无支持终止,满足非劣效性要求,可实现成功恢复。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16834 2026-08-18 cs.CL cs.AI 新提交 73%

Model Hypnosis: Strong control of AI via additive subliminal effects

模型催眠:通过叠加阈下效应对AI进行强控制

Enric Boix-Adsera, Benedict Tessler

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 该研究发现AI模型普遍存在模型催眠现象,可通过组合提示中微弱无关线索强控制模型行为,该现象跨模型家族且具迁移性,对AI安全与可解释性构成挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13262 2026-08-14 cs.LG cs.AI 新提交 73%

Into the ORBIT for Time Series: Training Regimes for Foundation Models

面向时间序列的ORBIT:基础模型的训练机制

Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong

专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 该研究针对时间序列基础模型训练分布控制不足的问题,提出ORBIT训练范式,结合多级采样与增量训练,训练Falcon-2.0模型并引入Rank引导跨深度对齐,在多基准测试中展现优异零样本预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11181 2026-08-12 cs.CC cs.AI cs.LG 新提交 73%

How to Verify Consistency of Probabilistic Claims

如何验证概率断言的一致性

Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 该研究针对概率预测器答案的自洽性验证问题,基于Nilsson的工作构造交互式PCP,为概率预测器自洽性证明提供复杂性理论基础,是训练模型证明自身一致性的第一步。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06876 2026-08-10 cs.CV cs.AI cs.DC cs.LG 新提交 73%

FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition

FedVAR:用于视频异常识别的原型对齐联邦框架

Ghani Haider, Majid Kundroo, Boyun Eom, Dong Hwan Park, Chen Chen, Taehong Kim

机构 * Chungbuk National University(忠北国立大学) Electronics and Telecommunications Research Institute (ETRI)(电子通信研究院) University of Central Florida(中佛罗里达大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 针对联邦视频异常识别中的语义错位问题,本文提出FedVAR框架,通过视觉语言模型的原型对齐机制缓解错位,实验显示其性能优于现有联邦基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02820 2026-08-05 cs.CR cs.AI cs.LG 新提交 73%

Evading Chain-of-Thought Monitoring Through Model Poisoning

通过模型投毒规避思维链监控

Giorgio Severi, Shujaat Mirza, Blake Bullwinkel, Amanda Minnich

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 该研究从模型投毒视角,通过微调或课程训练向推理模型植入CoT隐藏后门,使其产生攻击者选定行为的同时隐藏轨迹,揭示CoT监控应关注轨迹与响应的一致性而非轨迹内部异常。

Comments 15 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25633 2026-07-29 cs.CL cs.AI 新提交 73%

Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

构造驱动注入:用于大语言模型的基于语言基础编辑的代码混合指纹

Yongyi Cui, Yue Li, Tianbao Jiang, Xin Yi

专题命中 其他安全 :alignment(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

AI总结 研究针对大语言模型易被滥用问题,提出统一指纹框架。通过LCF按规则构造代码混合指纹,再用LCFEdit结合多语言表示和跨语言对齐注入指纹,实现构造感知注入,确保更新稳定,能持续验证所有权且对模型效用影响小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29604 2026-06-30 cs.LG cs.AI 73%

Mechanistically Eliciting Latent Behaviors in Language Models

机械性地引发语言模型中的潜在行为

Andrew Mack, Nina Panickssery, Alexander Matt Turner

机构 * Principles of Intelligence(智能原理研究所) Anthropic(Anthropic公司) Independent(独立)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 提出因果扰动引发(CPE)方法,通过无监督方式发现可解释的低秩适配器,以揭示语言模型中的隐藏行为模式,并展示其在数据效率、安全评估和模型对齐方面的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22769 2026-06-23 cs.LG cs.AI 新提交 73%

Noise is Signal: Density-Based Outliers as Leading Indicators of Occupational Emergence in Labor Market Text

噪声即信号:基于密度的异常值作为劳动力市场文本中职业涌现的领先指标

Shreyash Rawat

机构 * Independent Researcher(独立研究员)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出涌现-密度反转假设,证明低密度异常值预示新职业涌现,通过时序分析验证并改进涌现职业评分,预测准确率从F1=0.61提升至0.74。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12913 2026-06-15 cs.AI cs.LG cs.NE 版本更新 73%

Actionable Interpretability Must Be Defined in Terms of Symmetries

可操作的可解释性必须根据对称性来定义

Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra

机构 * University of Oxford(牛津大学) ETH Zurich(苏黎世联邦理工学院) University of Cambridge(剑桥大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文论证AI可解释性研究存在根本性问题,提出可操作的可解释性应基于四种对称性来定义,以形式化可解释模型并统一可解释推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05220 2026-06-08 cs.LG cs.AI 版本更新 73%

MidSteer: Optimal Affine Framework for Steering Generative Models

MidSteer:用于引导生成模型的最优仿射框架

Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu, Martin Benning, Gregory Slabaugh, Jiankang Deng, Ismail Elezi

机构 * University of Basel(巴塞尔大学) University of California, Berkeley(加州大学伯克利分校) ETH Zurich(苏黎世联邦理工学院) University of Cambridge(剑桥大学) University of Washington(华盛顿大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MidSteer,一种基于仿射变换的最优概念引导框架,通过最小干扰实现生成模型中的概念切换,并在视觉扩散模型和大型语言模型上验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02983 2026-06-02 cs.CL cs.AI 73%

Truth, Trust, and Trouble: Medical AI on the Edge

真相、信任与麻烦:边缘上的医疗AI

Mohammad Anas Azeez, Rafiq Ali, Ebad Shabbir, Zohaib Hasan Siddiqui, Gautam Siddharth Kashyap, Jiechao Gao, Usman Naseem

机构 * Jamia Hamdard(贾迈亚哈姆达德大学) DSEU-Okhla Macquarie University(麦考瑞大学) Center for SDGC, Stanford University(SDGC中心,斯坦福大学)

专题命中 其他安全 :safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

AI总结 通过一个包含1000多个健康问题的基准测试框架,评估Mistral-7B、BioMistral-7B-DARE和AlpaCare-13B三个模型在诚实、有用性和无害性方面的表现,发现AlpaCare-13B准确率最高(91.7%)且无害性最佳(0.92),而领域微调可提升安全性,少样本提示能提高准确率,但复杂查询下有用性下降。

Comments Accepted at EMNLP 2025 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29126 2026-05-29 cs.LG cs.AI 73%

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

何时与多久?时间推理中的读出-中介角度

Shreyas Fadnavis, Praitayini Kanakaraj, Felix Wyss

机构 * Bioscope AI

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 通过测量线性探针与模型实际计算子空间之间的角度,发现探针可能学习与模型无关的正交方向,从而揭示基于探针的可解释性存在根本缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏