arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8034 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8034 篇

2403.13812 2024-03-22 cs.DL cs.AI cs.CL cs.CY cs.LG stat.OT 70%

Quantitative Analysis of AI-Generated Texts in Academic Research: A Study of AI Presence in Arxiv Submissions using AI Detection Tool

Arslan Akram

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 8 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12893 2023-12-29 cs.RO cs.AI cs.CV 70%

A Safer Vision-based Autonomous Planning System for Quadrotor UAVs with Dynamic Obstacle Trajectory Prediction and Its Application with LLMs

Jiageng Zhong, Ming Li, Yinliang Chen, Zihang Wei, Fan Yang, Haoran Shen

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18762 2023-10-31 cs.LG cs.CR 70%

Purify++: Improving Diffusion-Purification with Advanced Diffusion Models and Control of Randomness

Boya Zhang, Weijian Luo, Zhihua Zhang

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01508 2023-10-10 cs.LG cs.CR cs.CV 70%

Circumventing Concept Erasure Methods For Text-to-Image Generative Models

Minh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal, Chinmay Hegde

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12833 2023-08-25 cs.CL cs.CR 70%

Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities

Maximilian Mozes, Xuanli He, Bennett Kleinberg, Lewis D. Griffin

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10970 2023-08-03 cs.LG 70%

Can GPT-4 Perform Neural Architecture Search?

Mingkai Zheng, Xiu Su, Shan You, Fei Wang, Chen Qian, Chang Xu, Samuel Albanie

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14784 2023-05-25 cs.AI cs.CL cs.CY cs.LG 70%

Anthropomorphization of AI: Opportunities and Risks

Ameet Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, Ashwin Kalyan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01299 2023-05-03 cs.LG cs.SY eess.SY 70%

An Improved Yaw Control Algorithm for Wind Turbines via Reinforcement Learning

Alban Puech, Jesse Read

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

Journal ref Amini, MR., Canu, S., Fischer, A., Guns, T., Kralj Novak, P., Tsoumakas, G. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2022. Lecture Notes in Computer Science(), vol 13717. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10513 2022-11-23 cs.AI 70%

Computable Artificial General Intelligence

Michael Timothy Bennett

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments Experiment code available on TechRxiv: https://www.techrxiv.org/articles/preprint/Computable_Artificial_General_Intelligence/19740190

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.12016 2021-08-30 cs.LG 70%

DeepFlow: Abnormal Traffic Flow Detection Using Siamese Networks

Sepehr Sabour, Sanjeev Rao, Majid Ghaderi

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

Comments 7 pages, 12 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11872 2021-06-23 cs.LG cs.NE 70%

Randomness In Neural Network Training: Characterizing The Impact of Tooling

Donglin Zhuang, Xingyao Zhang, Shuaiwen Leon Song, Sara Hooker

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments 21 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.10247 2021-01-26 cs.LG 70%

Incorporating Expert Guidance in Epidemic Forecasting

Alexander Rodríguez, Bijaya Adhikari, Naren Ramakrishnan, B. Aditya Prakash

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments Appears in SIGKDD 2020 epiDAMIK

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05651 2020-07-16 cs.LG stat.ML 70%

Bayesian Variational Autoencoders for Unsupervised Out-of-Distribution Detection

Erik Daxberger, José Miguel Hernández-Lobato

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments 21 pages, extended version with supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.04877 2018-08-29 cs.LG cs.CV cs.HC 70%

Learning via social awareness: Improving a deep generative sketching model with facial feedback

Natasha Jaques, Jennifer McCleary, Jesse Engel, David Ha, Fred Bertsch, Rosalind Picard, Douglas Eck

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.08476 2017-07-27 cs.AI cs.CR 70%

Guidelines for Artificial Intelligence Containment

James Babcock, Janos Kramar, Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06039 2026-07-08 cs.SE 新提交 69%

Automating Quality Assessment with NLP of LLM-Generated Defeaters

利用自然语言处理对大语言模型生成的反驳进行质量评估自动化

Tihomir Rohlinger, Daniel Ratiu, Stefan Wagner

专题命中 其他安全 :safety(abstract,journal_ref);alignment(abstract)

AI总结 研究针对大语言模型生成的反驳质量评估依赖人工且主观的问题,提出结合保证案例图结构特征、语义嵌入和元分类器的自动化评估方法,经案例研究验证,该方法能减少主观差异,为保证案例审查提供决策支持。

Comments 10 pages, 2 figures. Author preprint version of a paper published at ICSRS 2025

Journal ref 2025 9th International Conference on System Reliability and Safety (ICSRS), Turin, Italy, 2025, pp. 101 110

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23670 2026-08-26 cs.AI cs.CL cs.LG 新提交 67%

Automata from Agent Traces: Failure and Next-Step Prediction

基于智能体轨迹的自动机:失败与下一步预测

Seonglae Cho, Franklin Cardenoso Fernandez, Umar Mohammed, Zekun Wu, Kleyton Da Costa, Ilham Wicaksono, Adriano Koshiyama

机构 * Holistic AI(整体人工智能公司) PUC-Rio(里约热内卢天主教大学) University College London(伦敦大学学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出将LLM智能体的轨迹语料库压缩为有限状态机(FSM),该模型在12个公开数据集上表现紧凑、准确且构建快速,可同时提升下一步预测与失败预测性能,为安全监控提供模型无关的结构基元。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03058 2026-08-25 cs.CL cs.AI cs.CY 版本更新 67%

Verbalizing LLMs' assumptions to explain and control sycophancy

使LLM的假设 verbal化以解释和控制谄媚行为

Myra Cheng, Isabel Sieh, Humishka Zope, Sunny Yu, Lujain Ibrahim, Aryaman Arora, Jared Moore, Desmond Ong, Dan Jurafsky, Diyi Yang

机构 * Stanford University(斯坦福大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出Verbalized Assumptions框架,通过揭示LLM对用户的错误假设来解释和控制其谄媚行为,发现LLM在面对用户提问时倾向于提供验证而非客观信息,进而提出可解释的细粒度控制方法。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17850 2026-08-24 cs.HC cs.AI cs.CL cs.CY 版本更新 67%

Mind the Style: Impact of Communication Style on Human-Chatbot Interaction

注意风格:沟通风格对人机交互的影响

Erik Derner, Dalibor Kučera, Aditya Gulati, Ayoub Bagheri, Nuria Oliver

机构 * Czech Technical University in Prague(捷克技术大学(布拉格)) University of South Bohemia(南波西米亚大学) Utrecht University(乌得勒支大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究探讨了沟通风格对人机交互的影响,发现友好风格提升女性用户满意度和任务完成率,且用户较少模仿聊天机器人风格。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18339 2026-08-20 cs.CV cs.AI cs.CL cs.LG 新提交 67%

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

从推理到适应:视觉语言模型的统一最优传输视角

Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊公司) Wayne State University(韦恩州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出名为algname的VLM测试时适应方法,通过Wasserstein最优传输统一VLM的推理与适应目标,实验显示其性能较最优方法提升7%且效率先进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15794 2026-08-19 cs.LG cs.AI cs.CL 版本更新 67%

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

自蒸馏作为大语言模型的性能恢复机制:对抗压缩与灾难性遗忘

Chi Liu, Xin Chen, Xu Zhou, Fangbo Tu, Srinivasan Manoharan

机构 * PayPal AI

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于自蒸馏微调的性能恢复框架,通过理论分析和实验验证,证明自蒸馏能有效恢复模型能力,揭示了高维流形对齐与性能恢复之间的强相关性。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06416 2026-08-19 cs.CL cs.AI cs.LG 版本更新 67%

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

注意力流动:通过故事摘要追踪LLM概念性参与

Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan

机构 * Cornell University(康奈尔大学) Aarhus University(奥胡斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过比较人类与LLM生成的故事摘要,分析模型在文本中的概念性参与模式,发现模型更关注文本结尾,揭示了摘要生成任务的复杂性。

Comments Error found in data creation pipeline

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13925 2026-08-17 cs.LG cs.AI cs.CL 新提交 67%

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

CForce:通过一致性强制提升扩散大语言模型(dLLMs)的并行解码性能

Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CForce方法,通过一致性强制提升dLLMs的并行解码性能,在LLaDA模型上实验证实其在高并行解码预算下可优化速度-质量权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10703 2026-08-12 cs.LG cs.AI cs.CL cs.HC 新提交 67%

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

你的大语言模型,你的风格:用于大语言模型行为控制的行为模式轴

Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出情境化行为数据框架,构建3200个对比性行为场景,发现LLMs有稳定且模型特异性的行为特征,提出行为模式轴控制LLM行为,表明其类人格倾向是可测量可控的行为模式。

Comments 33 pages, 8 figures. Code and data: https://github.com/lhz191/LLM-Behavioral-Personality

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06578 2026-08-12 cs.AI cs.CL cs.LG 交叉投稿 67%

Divergent Response Modes in Frontier Language Models Under Steering Pressure

前沿语言模型在引导压力下的发散响应模式

Ali Jalal-Kamali

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究评估六种前沿语言模型在引导压力下的响应差异,发现模型间存在不同响应模式,以Llama为对象追溯到内部机制,相关发现经多实验验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07881 2026-08-11 cs.AI cs.CL cs.IT cs.LG math.IT 新提交 67%

GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering

GRACE:用于可扩展混合数据聚类的大语言模型(LLM)基础语义度量空间

Zihua Yang, Zhencheng Xie, Junyang Chen, Liang Xie, Yiqun Zhang, Mengke Li, Yang Lu

机构 * Guangdong University of Technology(广东工业大学) Tsinghua University(清华大学) Peking University(北京大学) Shenzhen University(深圳大学) Xiamen University(厦门大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对混合数据聚类中语义丰富度与可扩展性的矛盾,提出GRACE框架,通过多视角LLM查询实现一次性语义嵌入,兼顾可扩展性与聚类性能,优于11种对比方法。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14888 2026-08-06 cs.CL cs.AI cs.CV cs.LG 版本更新 67%

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

推理动态与监控模态依赖性的局限性

Danae Sánchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras, Desmond Elliott

机构 * University of Copenhagen, Department of Computer Science(哥本哈根大学计算机科学系) University of Sheffield, School of Computer Science(谢菲尔德大学计算机科学学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究分析18种VLM的推理动态,发现模型易受早期预测影响,推理训练模型在模态条件下表现出更强的纠正行为,但其效果依赖于模态条件,且CoT的可检测性因模型而异。

Comments Accepted for publication in COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01868 2026-08-04 cs.CY cs.AI cs.CL cs.MA 新提交 67%

No One Wins in Nuclear War: A Social Simulation of Military Decision-making

核战争中无人取胜:军事决策的社会模拟

Glenn Matlin, Isaac Song, Anthony Wen-Ming Zang, Mark Riedl

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究构建了名为WOPR的社会模拟环境,以兵棋推演为载体,采用Concordia为默认工具,实现了带有可验证规则引擎和四级沟通阶梯的模拟系统,将军事决策等战略选择转化为智能体的明确决策。

Comments 16 pages, 11 figures. Published at the Social Sim'26 Workshop at COLM 2026. Code and replay data: https://github.com/eilab-gt/wopr

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00502 2026-08-04 cs.CV 新提交 67%

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

SpatialAfford:教紧凑视觉语言模型(VLMs)关注何处与定位何处以实现 affordance

Yufei Zhang, Chenlu Zhan, Donghui Sun, Xiaoxin Chen, Hongwei Wang

机构 * Zhejiang University(浙江大学) Vivo Mobile Communication Co., Ltd(维沃移动通信有限公司)

专题命中 其他安全 :alignment(abstract,abstract_cn)

AI总结 SpatialAfford 是一个两阶段框架,通过空间注意力对齐和空间感知 GRPO,使紧凑 4B VLMs 在多个基准上的 affordance 定位性能优于更强的 7B+ 基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00086 2026-08-04 cs.CV cs.AI cs.CL cs.LG 版本更新 67%

Hierarchical Pre-Training of Vision Encoders with Large Language Model

基于大语言模型的视觉编码器分层预训练

Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee

机构 * University of Cincinnati(辛辛那提大学) National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出HIVE框架,通过引入视觉编码器与大语言模型间的分层交叉注意力机制,提升视觉语言对齐,改进特征融合与表征学习,实验表明其在图像分类和多模态任务中表现优异。

Comments 17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏