arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8033 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8033 篇

2411.05040 2024-11-11 cs.CL cs.AI 73%

Bottom-Up and Top-Down Analysis of Values, Agendas, and Observations in Corpora and LLMs

Scott E. Friedman, Noam Benkler, Drisana Mosaphir, Jeffrey Rye, Sonja M. Schmer-Galunder, Micah Goldwater, Matthew McLure, Ruta Wheelock, Jeremy Gottlieb, Robert P. Goldman, Christopher Miller

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13697 2024-09-24 cs.CL cs.AI 73%

Prompt Baking

Aman Bhargava, Cameron Witkowski, Alexander Detkov, Matt Thomson

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19348 2024-08-09 cs.LG cs.AI 73%

Deep Learning for Cross-Domain Data Fusion in Urban Computing: Taxonomy, Advances, and Outlook

Xingchen Zou, Yibo Yan, Xixuan Hao, Yuehong Hu, Haomin Wen, Erdong Liu, Junbo Zhang, Yong Li, Tianrui Li, Yu Zheng, Yuxuan Liang

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Journal ref Inform.Fusion.113(2025)102606

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19334 2024-06-11 cs.AI cs.CL cs.CV cs.MM cs.SD 73%

LLMs Meet Multimodal Generation and Editing: A Survey

Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian, Hongyu Liu, Xiaowei Chi, Runtao Liu, Ruibin Yuan, Yazhou Xing, Wenhai Wang, Jifeng Dai, Yong Zhang, Wei Xue, Qifeng Liu, Yike Guo, Qifeng Chen

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 52 Pages with 16 Figures, 12 Tables, and 545 References. GitHub Repository at: https://github.com/YingqingHe/Awesome-LLMs-meet-Multimodal-Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12582 2024-03-13 cs.CY cs.AI 73%

Understanding and Avoiding AI Failures: A Practical Guide

Heather M. Williams, Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03096 2024-02-14 cs.LG cs.AI cs.NE 73%

What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes

Victor Lecomte, Kushal Thaman, Rylan Schaeffer, Naomi Bashkansky, Trevor Chow, Sanmi Koyejo

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00323 2024-01-19 cs.AI cs.LG 73%

Thought Cloning: Learning to Think while Acting by Imitating Human Thinking

Shengran Hu, Jeff Clune

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2023 as a spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00667 2023-09-06 cs.CL cs.LG 73%

Taken out of context: On measuring situational awareness in LLMs

Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, Owain Evans

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.02617 2021-06-07 cs.AI cs.LG 73%

Be Considerate: Objectives, Side Effects, and Deciding How to Act

Parand Alizadeh Alamdari, Toryn Q. Klassen, Rodrigo Toro Icarte, Sheila A. McIlraith

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.07421 2021-01-05 cs.CV cs.AI cs.LG cs.MM 73%

Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models

Tong Che, Xiaofeng Liu, Site Li, Yubin Ge, Ruixiang Zhang, Caiming Xiong, Yoshua Bengio

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04071 2020-08-11 cs.CY cs.AI 73%

On Controllability of AI

Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.06497 2019-12-16 cs.CR cs.AI cs.LG 73%

Founding The Domain of AI Forensics

Ibrahim Baggili, Vahid Behzadan

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for presentation at SafeAI2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.05590 2018-11-15 cs.LG cs.AI stat.ML 73%

Emergence of Addictive Behaviors in Reinforcement Learning Agents

Vahid Behzadan, Roman V. Yampolskiy, Arslan Munir

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23921 2026-08-26 cs.CV 新提交 71%

HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment

HAP:基于跨模态对齐的头部自适应视觉Token剪枝

Yuanhao Sun, Huawei Ji, Yuan Jin, Cheng Deng, Luoyi Fu, Xinbing Wang

机构 * Shanghai Jiao Tong University(上海交通大学) University of Edinburgh(爱丁堡大学)

专题命中 其他安全 :alignment(title)

AI总结 该研究针对视觉-语言模型预填充成本过高的问题,提出基于PAQ指标的头部自适应视觉Token剪枝方法,在18个基准上实现最优权衡,在LLaVA-1.5-7B上仅保留5.6% Token即可维持99.1%的原始性能,优于基线AutoPrune。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20969 2026-08-24 cs.CV 新提交 71%

Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis

用于模式对齐的运动学知识图谱:多模态步态分析中的结构化潜在表示学习

Chen Dong, He Zonglin, Cheung Kenneth M. C

专题命中 其他安全 :alignment(title)

AI总结 本研究提出ScoliDetect框架,通过运动学知识图谱(KKM)实现多模态步态分析,其介导的融合结合三模态对比预训练提升了脊柱侧凸筛查的泛化性与可解释性,外部ROC-AUC达0.972。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16041 2026-08-18 cs.RO 新提交 71%

ScenarioCharacterization: A Modular Toolkit for Characterizing Safety across Trajectory Datasets

ScenarioCharacterization:用于刻画轨迹数据集安全性的模块化工具包

Ingrid Navarro, Yutong Duan, Jonathan Francis, Jean Oh

机构 * Robotics Institute, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机学院机器人研究所) Stack AV Bosch Center for Artificial Intelligence(博世人工智能中心)

专题命中 其他安全 :safety(title)

AI总结 本文提出开源模块化框架ScenarioCharacterization,可自动刻画轨迹数据集的驾驶场景安全性,适配多数据集,在Waymo等数据集上验证,支持下游应用。

Comments 9 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20936 2026-08-14 physics.soc-ph nlin.AO 版本更新 71%

Empathy Modeling in Active Inference Agents for Perspective-Taking and Alignment

主动推断代理中的共情建模:用于视角转换与对齐

Mahault Albarracin, Hongju Pae, Philip Wilson, Anna Mikeda, Alejandro Jimenez-Rodriguez, Sanjeev V. Namjoshi, Harshil Shah

专题命中 其他安全 :alignment(title)

AI总结 该研究提出了一种基于主动推断的共情框架,通过视角转换实现稳定合作,揭示了共情结构对社会协调的关键作用。

Comments Code and data: https://doi.org/10.5281/zenodo.21908008

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20318 2026-08-04 cs.CV cs.MM 版本更新 71%

UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

UniCVR:从对齐到重排序的统一零样本复合视觉检索

Haokun Wen, Xuemeng Song, Haoyu Zhang, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学) Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室)

专题命中 其他安全 :alignment(title)

AI总结 UniCVR提出首个统一零样本复合视觉检索框架,结合多模态大语言模型与视觉语言预训练模型,通过两阶段方法实现多任务联合优化,实验表明其在五个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23523 2026-07-28 cs.AR 新提交 71%

CircuitWeave: Topology-Behavior Alignment for Executable Multimodal RTL Generation

CircuitWeave:用于可执行多模态RTL生成的拓扑-行为对齐

Jiahao Feng, Haiyan Qin, Zhiwei Xie, Wang Kang

专题命中 其他安全 :alignment(title)

AI总结 研究如何从自然语言规范生成RTL,提出CircuitWeave合同介导多模态框架,从原理图和文本提取合同并融合,通过联合目标监督相关过程,在生成可执行RTL上取得较好效果,部分指标高于无原理图的情况。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19081 2026-06-18 q-bio.NC cs.HC 新提交 71%

Retrieval-Based Brain Decoding by Alignment, not Complexity

基于对齐而非复杂性的检索式脑解码

Matteo Ciferri, Matteo Ferrante, Nicola Toschi

专题命中 其他安全 :alignment(title)

AI总结 本文通过跨多数据集实验证明,线性对比解码器在脑解码中优于岭回归和标准非线性方法,表明解码增益更多来自训练目标而非架构复杂性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18441 2026-06-18 cs.CV 新提交 71%

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs

推理即交集:视频多模态大语言模型中视觉焦点的一致性帧对齐

Chengwen Liu, Zhe Huang, Jisheng Dang, Hong Peng, Qi Tian, Tat-Seng Chua

机构 * School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) Beijing University of Posts and Telecommunications(北京邮电大学) Cloud and AI BU, Huawei(华为云与AI业务部) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 其他安全 :alignment(title)

AI总结 提出无时间标注的过程级奖励框架CF-GRPO,通过视频内在线索构建一致性帧先验,并利用一致性帧奖励优化模型帧使用与先验的对齐,提升视频推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13190 2026-06-12 cs.RO cs.HC 新提交 71%

Multi-Modal Multi-Agent Robotic Cognitive Alignment enabled by Non-Invasive Consumer Brain Computer Interfaces: A Proof of Concept Exploration

基于非侵入式消费级脑机接口的多模态多智能体机器人认知对齐:概念验证探索

Nataliya Kosmyna, Liz Jenkins, Anoop K. Sinha

机构 * GOOGLE(谷歌) Paradigms of Intelligence(智能范式) Cambridge, MA, United States(马萨诸塞州剑桥市,美国) Mountain View, CA, United States(加利福尼亚州山景城,美国)

专题命中 其他安全 :alignment(title)

AI总结 提出一种框架,利用消费级脑机接口监测脑电信号,在高认知负荷时延迟智能体通信,实现认知对齐的多智能体交互,初步验证了实时信号处理、大语言模型与机器人结合的可行性。

Comments 19 pages, 9 figures, for associated video, see https://youtu.be/0Tav-G87XGs

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.01824 2026-06-04 eess.SY cs.SY 71%

Robust Distance-Based Formation Control of Multiple Rigid Bodies with Orientation Alignment

多刚体鲁棒距离基形成控制

Alexandros Nikou, Christos K. Verginis, Dimos V. Dimarogonas

专题命中 其他安全 :alignment(title)

AI总结 本文研究了在3D空间中,针对第二类非线性多智能体系统,设计了去中心化无模型控制协议,实现距离和方向基的形成控制,并通过仿真验证了控制器性能。

Comments IFAC Word Congress 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01970 2026-06-02 cs.RO cs.MA cs.SY eess.SY 71%

Market-Based Replanning for Safety-Critical UAV Swarms in Search and Rescue Missions

基于市场重规划的搜救任务中安全关键无人机群

Luiz Giacomossi, Andrea Haglund, Claire Namatovu, Emily Zainali, Esaias Målqvist, Yonatan M. Beyene, Ivan Tomasic, Baran Çürüklü, Håkan Forsberg

机构 * KTH Royal Institute of Technology(皇家理工学院) Swedish Defence Research Agency(瑞典国防研究机构) KTH Royal Institute technological Institute(皇家理工学院)

专题命中 其他安全 :safety(title)

AI总结 提出一种分布式协调架构IRDS,通过反向拍卖市场机制和几何共识协议,在无人机故障下自主重分配任务,在25%退化下保持93%任务成功率。

Comments 6 pages, 4 figures, accepted at MIPRO 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01562 2026-05-22 quant-ph 71%

Kernel Alignment for Quantum Support Vector Machines Using Genetic Algorithms

使用遗传算法的量子支持向量机核的核对齐

Floyd M. Creevey, Jamie A. Heredge, Martin E. Sevior, Lloyd C. L. Hollenberg

专题命中 其他安全 :alignment(title)

AI总结 本文提出了一种基于遗传算法的量子支持向量机核对齐方法,通过评估监督和无监督核损失函数对编码电路优化的影响,提高了分类准确率,并在金融、医疗和材料科学等领域展示了改进的机器学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00429 2026-04-14 eess.SY cs.SY 71%

Distributed Safety-Critical Control of Multi-Agent Systems with Time-Varying Communication Topologies

多智能体系统中具有时间变化通信拓扑的分布式安全控制

Shiyu Cheng, Luyao Niu, Bhaskar Ramasubramanian, Andrew Clark, Radha Poovendran

专题命中 其他安全 :safety(title)

AI总结 本文提出一种基于分布式优化的控制框架,通过截断函数和辅助不匹配变量处理时间变化通信拓扑下的多智能体协同与避障问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25460 2026-03-27 cs.SD 71%

CLAR: CIF-Localized Alignment for Retrieval-Augmented Speech LLM-Based Contextual ASR

CLAR:基于检索增强的语音大语言模型的上下文语音识别中的CIF局部对齐

Shangkun Huang, Huan Shen, Wei Zou, Yunzhang Chen

机构 * BRVoice Team, Bairong, Inc., China(BRVoice团队,百融云创,中国)

专题命中 其他安全 :alignment(title)

AI总结 CLAR通过CIF学习单调性token级对齐,提升语音识别中专有名词和长尾词的识别效果,减少表示稀释和注意力漂移,提高检索准确率。

Comments Submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01477 2026-03-03 cs.RO 71%

SFCo-Nav: Efficient Zero-Shot Visual Language Navigation via Collaboration of Slow LLM and Fast Attributed Graph Alignment

SFCo-Nav: 通过慢LLM与快速属性图对齐的协作实现高效的零样本视觉语言导航

Chaoran Xiong, Litao Wei, Xinhao Hu, Kehui Ma, Ziyi Xia, Zixin Jiang, Zhen Sun, Ling Pei

机构 * Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai Jiao Tong University(上海导航与位置基于服务重点实验室,上海交通大学)

专题命中 其他安全 :alignment(title)

AI总结 SFCo-Nav通过慢LLM与快速属性图对齐的协作,实现了高效的零样本视觉语言导航,显著提升效率并降低计算成本。

Comments Accepted by 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23068 2026-02-27 cs.SD 71%

TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment

TADA: 一种通过文本-语音双模对齐生成语音模型的框架

Trung Dang, Sharath Rao, Ananya Gupta, Christopher Gagne, Panagiotis Tzirakis, Alice Baird, Jakub Piotr Cłapa, Peter Chin, Alan Cowen

机构 * Hume AI Dartmouth College(达特茅斯学院)

专题命中 其他安全 :alignment(title)

AI总结 TADA通过文本-语音双模对齐生成框架,实现语音与文本的一对一同步建模,提升语音生成的保真度和效率,减少幻觉并降低推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10593 2026-02-12 cs.CV 71%

Fast Person Detection Using YOLOX With AI Accelerator For Train Station Safety

利用YOLOX与AI加速器的快速人员检测用于车站安全

Mas Nurul Achmadiah, Novendra Setyawan, Achmad Arif Bryantono, Chi-Chia Sun, Wen-Kai Kuo

机构 * Department of Electro-Optics, National Formosa University, Taiwan(国立Formosa大学电子光学系) Department of Electrical Engineering, National Taipei University, Taiwan(国立台北大学电子工程系) Department of Electrical Engineering, University of Muhammadiyah Malang, Indonesia(穆罕默迪亚大学Malang分校电子工程系) Department of Electronics Engineering, State Polytechnic of Malang, Indonesia(Malang州立理工学院电子工程系) Smart Manufacturing and Intelligent Machinery Research Center, National Formosa University, Taiwan(国立Formosa大学智能制造与智能机械研究中心) Department of Electronics Engineering, National Formosa University, Taiwan(国立Formosa大学电子工程系)

专题命中 其他安全 :safety(title)

AI总结 本文提出利用YOLOX与Hailo-8 AI加速器提高车站乘客检测的准确性和效率

Comments 6 pages, 8 figures, 2 tables. Presented at 2024 International Electronics Symposium (IES). IEEE DOI: 10.1109/IES63037.2024.10665874

Journal ref 2024 International Electronics Symposium (IES), pp. 504-509, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏