arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8034 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8034 篇

2404.05103 2024-04-09 cs.HC 71%

Chart What I Say: Exploring Cross-Modality Prompt Alignment in AI-Assisted Chart Authoring

Nazar Ponochevnyi, Anastasia Kuzminykh

专题命中 其他安全 :alignment(title)

Comments Will be published In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10958 2024-01-23 eess.IV 71%

Detection of Thermal Events by Semi-Supervised Learning for Tokamak First Wall Safety

Christian Staron, Hervé Le Borgne, Raphaël Mitteau, Erwan Grelier, Nicolas Allezard

专题命中 其他安全 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08535 2023-11-16 cs.DB 71%

Taxonomy, Semantic Data Schema, and Schema Alignment for Open Data in Urban Building Energy Modeling

Liang Zhang, Jianli Chen, Jia Zou

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01317 2023-09-11 cs.CV eess.IV 71%

ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders

Shawn Xu, Lin Yang, Christopher Kelly, Marcin Sieniek, Timo Kohlberger, Martin Ma, Wei-Hung Weng, Atilla Kiraly, Sahar Kazemzadeh, Zakkai Melamed, Jungyeon Park, Patricia Strachan, Yun Liu, Chuck Lau, Preeti Singh, Christina Chen, Mozziyar Etemadi, Sreenivasa Raju Kalidindi, Yossi Matias, Katherine Chou, Greg S. Corrado, Shravya Shetty, Daniel Tse, Shruthi Prabhakara, Daniel Golden, Rory Pilgrim, Krish Eswaran, Andrew Sellergren

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07275 2022-11-15 cs.CV 71%

Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment

Junyang Wang, Yi Zhang, Ming Yan, Ji Zhang, Jitao Sang

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.02830 2021-06-08 eess.AS 71%

Reinforce-Aligner: Reinforcement Alignment Search for Robust End-to-End Text-to-Speech

Hyunseung Chung, Sang-Hoon Lee, Seong-Whan Lee

专题命中 其他安全 :alignment(title)

Comments Accepted in INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12651 2020-05-19 cs.PL 71%

Type Safety with JSON Subschema

Andrew Habib, Avraham Shinnar, Martin Hirzel, Michael Pradel

专题命中 其他安全 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13900 2020-04-01 cs.IR cs.SI 71%

A large-scale Twitter dataset for drug safety applications mined from publicly existing resources

Ramya Tekumalla, Juan M. Banda

专题命中 其他安全 :safety(title)

Comments 8 tables, 2 figures, 7 pages, accepted after peer review as a workshop paper in ACM Conference on Health, Inference, and Learning (CHIL) 2020 https://www.chilconference.org/agenda/

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.11861 2020-02-28 cs.MA eess.SP 71%

Simulation of Real-time Routing for UAS traffic Management with Communication and Airspace Safety Considerations

Zhao Jin, Ziyi Zhao, Chen Luo, Franco Basti, Adrian Solomon, M. Cenk Gursoy, Carlos Caicedo, Qinru Qiu

专题命中 其他安全 :safety(title)

Comments The 38th AIAA/IEEE Digital Avionics Systems Conference (DASC)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00520 2018-07-25 astro-ph.SR 71%

Tracking the spin axes orbital alignment in selected binary systems - Torun Rossiter-McLaughlin effect survey

P. Sybilski, R. K. Pawłaszek, A. Sybilska, M. Konacki, K. G. Hełminiak, S. K. Kozłowski, M. Ratajczak

专题命中 其他安全 :alignment(title)

Comments 30 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.04766 2017-08-17 cond-mat.mes-hall 71%

Fundamental Band Gap and Alignment of Two-Dimensional Semiconductors Explored by Machine Learning

Zhen Zhu, Baojuan Dong, Teng Yang, Zhi-Dong Zhang

专题命中 其他安全 :alignment(title)

Comments 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.02588 2017-01-11 q-bio.QM 71%

Synthesis of Methotrexate loaded Cerium fluoride nanoparticles with pH sensitive extended release coupled with Hyaluronic acid receptor with plausible theranostic capabilities for preclinical safety studies

Nitish Manu George

专题命中 其他安全 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24758 2026-08-26 cs.AI 新提交 70%

RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons

RACE:LLM神经元中功能一致性的可扩展统计估计

Runyu Wang, Bo Liu, Xiaxin Zhang, Yu Han, Jiawei Cao, Xiaoye Zhang, Zhe Zhang, Yifan Yang, Peng Ping

机构 * Nantong University(南通大学) Chongqing University of Post and Telecommunications(重庆邮电大学) China Southern Power Grid Company Limited(中国南方电网有限责任公司) Meituan(美团)

专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.AI

AI总结 针对LLM神经元功能一致性的可扩展统计估计难题,提出RACE框架,其领域特异性更优、计算开销低两个数量级,可有效评估Transformer神经元的全领域功能一致性。

Comments EMNLP-26 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11955 2026-08-14 cs.CY cs.HC 版本更新 70%

Philosophical vertigo with artificial intelligence

人工智能引发的哲学眩晕

Thomas A. Pollak, Hamilton Morrin, Murray Shanahan

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CY

AI总结 该研究提出人工智能引发的哲学眩晕概念,分析其产生、传播路径,关联临床妄想案例,指出AI将参与重构人类认知环境,并提出哲学可修正性作为应对对策。

Comments 29 pages, no figures. Source formatting revised to improve arXiv HTML accessibility; article text unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03538 2026-08-11 cs.LG 版本更新 70%

Online Learnability of Chain-of-Thought Verifiers: Soundness and Completeness Trade-offs

链式思维验证器的在线可学习性:正确性与完备性的权衡

Maria-Florina Balcan, Avrim Blum, Kiriaki Fragkia, Zhiyuan Li, Dravyansh Sharma

机构 * Carnegie Mellon University(卡内基梅隆大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) Northwestern University(西北大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本文提出一种在线学习框架,用于学习链式思维验证器,通过检查解决方案的正确性,解决生成器与验证器之间的反馈循环导致的分布偏移问题,并引入新的Littlestone维度扩展以优化验证器的学习。

Comments The abstract has been abridged due to arXiv length constraints

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02491 2026-08-06 cs.AI 版本更新 70%

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

长期测量:迈向对人机交互的纵向理解

Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel

机构 * Google Research(谷歌研究院) Cornell University(康奈尔大学) Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究针对语言模型融入生活引发的长期人机交互风险,结合社会科学测量与NLP计算方法,提出通过长期测量建模人类行为变化,实现问题行为在线检测以缓解用户长期风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13904 2026-07-24 cs.AI 版本更新 70%

Diagnosing Pathological Chain-of-Thought in Reasoning Models

诊断推理模型中的病理链式思维

Manqing Liu, David Williams-King, Ida Caspary, Linh Le, Hannes Whittingham, Puria Radmard, Cameron Tice, Edward James Young

机构 * Department of Epidemiology, CAUSALab, Harvard University, Boston, USA(流行病学系、CAUSALab、哈佛大学) Imperial College London, London, UK(伦敦帝国学院) McGill University, Montreal, Canada(麦吉尔大学) Geodesic Research, Cambridge, UK(Geodesic研究)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出了一种评估链式思维推理模型中病态的实用工具,通过定义具体度量标准和训练特定模型生物来识别和区分三种不同的病态现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31614 2026-07-01 eess.SY cs.AI cs.SY 新提交 70%

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

利用知识图谱和大语言模型自动化因果规范生成

Javal Vyas, Milapji Singh Gill, Mehmet Mercangöz

机构 * Autonomous Industrial Systems Lab, Imperial College London(帝国理工学院自主工业系统实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 提出一种语义AI框架,结合知识图谱与约束大语言模型,自动生成因果逻辑和操作安全叙述,减少手动工作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03335 2026-06-30 cs.CL 70%

Compressed Sensing for Capability Localization in Large Language Models

压缩感知在大语言模型能力定位中的应用

Anna Bair, Yixuan Even Xu, Mingjie Sun, J. Zico Kolter

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

AI总结 研究通过压缩感知方法识别大语言模型中特定能力依赖的稀疏注意力头,发现关闭少量头可显著降低特定能力表现,揭示了模型模块化组织原则。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23668 2026-06-23 cs.LG 新提交 70%

On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

关于提示条件语言模型作为通用学习器的局限性

David Mguni, Julian Ma, Jun Wang

机构 * Queen Mary University London(伦敦玛丽女王大学) University College London(伦敦大学学院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本文通过廉价谈话博弈模型分析提示条件语言模型,证明语言作为容量受限通道导致任务不可区分性,并因对齐约束产生不可约误差,从而否定其通过提示实现通用问题求解的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14580 2026-06-15 cs.CL 新提交 70%

Persuasion Index: A Theory-Guided Framework for Persuasion Analysis

说服指数:一个理论指导的说服分析框架

Liancheng Gong, Zhiyang Wang, Yiwei Xu, Julia Mendelsohn

机构 * University of Maryland, College Park(马里兰大学帕克分校) New York University(纽约大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

AI总结 提出基于心理学和传播学理论的15维说服指数(PI)及55个子特征实现,在四个数据集上验证其能解释说服相关修辞模式,并提供轻量级预测信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11399 2026-06-11 cs.CL 新提交 70%

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

基于场景的大型语言模型文化价值观探测与引导——扩展版

Trung Duc Anh Dang, Tung Kieu, Sarah Masud

机构 * University of Copenhagen(哥本哈根大学) Aalborg University(奥尔堡大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 提出基于场景的行为困境方法,通过令牌级概率和激活引导探测并调整LLM在英格尔哈特-韦尔泽尔文化轴上的潜在价值观,发现不同文化维度的引导存在耦合效应。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13053 2026-06-04 cs.AI cs.SY eess.SY 70%

Unpredictability of AI

AI的不可预测性

Roman V. Yampolskiy

机构 * Computer Engineering and Computer Science University of Louisville(计算机工程与计算机科学路易斯维尔大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文研究了AI安全领域中一个核心问题,即智能系统的行为预测难题,证明了即使知道终端目标,也无法准确预测超人类智能系统的行为,对AI安全产生了深远影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08426 2026-06-03 cs.GT cs.AI 70%

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

机制设计是不够的:面向合作AI的亲社会智能体

Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro, Van Q. Truong, Bernhard Schölkopf, Emanuele La Malfa, Zhijing Jin

机构 * ETH Zürich(苏黎世联邦理工学院) University of Oxford(牛津大学) Institute for Decentralized AI(去中心化人工智能研究所) Jinesis Lab, University of Toronto & Vector Institute(多伦多大学Jinesis实验室及向量研究所) EuroSafeAI University of Pennsylvania(宾夕法尼亚大学) Max Planck Institute for Intelligent Systems, Tübingen, Germany(德国图宾根最大计划智能系统研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文证明仅靠机制设计无法最大化LLM智能体的社会福利,并提出亲社会智能体(兼顾他人福利)能弥补这一差距,实现更优的社会与个体结果。

Comments 42 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28896 2026-05-29 cs.LG 70%

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models

LoRA适配器的特征几何:微调语言模型中表示差异的稀疏自编码器分析

Prasanth K K

机构 * Independent AI Safety Researcher(独立人工智能安全研究员)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本研究使用稀疏自编码器分析LoRA微调引起的表示几何变化,发现LoRA特征字典与预训练特征存在弱几何对齐,且适配器特定SAE能更有效重建delta激活。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21958 2026-05-22 cs.CL 70%

Diagnosis Is Not Prescription: Linguistic Co-Adaptation Explains Patching Hazards in LLM Pipelines

诊断并非处方:语言共适应解释了LLM流水线中的修补危害

Yoon Jeonghun, Kim Dongchan

机构 * KAIST (Korea Advanced Institute of Science and Technology)(韩国科学技术院) NAVER Corp.(NAVER公司)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 本文研究了多模块LLM代理失败时,诊断与修补之间的矛盾,发现路由模块虽为瓶颈,但注入修正示例反而降低性能,而修正查询重写模块则更有效,提出了语言合同假说解释这种现象。

Comments Preprint. Under review at EMNLP 2026 (ARR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21539 2026-05-22 cs.LG 70%

DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models

DualOptim+: 联合与解耦优化器状态的桥梁以提升大语言模型中的机器反遗忘

Xuyang Zhong, Qizhang Li, Yiwen Guo, Chen Liu

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Independent Researcher(独立研究者)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本文提出DualOptim+,一种改进大语言模型中机器反遗忘的新优化框架,通过引入基础状态和delta状态,有效平衡遗忘与保留目标,同时提出8位量化变体以减少内存开销,实验表明其在多个任务中均表现出色。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02193 2026-05-22 cs.AI 70%

From monoliths to modules: Decomposing transducers for efficient world modelling

从整体到模块:分解转换器以实现高效的world建模

Alexander Boyd, Franz Nowak, David Hyland, Manuel Baltieri, Fernando E. Rosas

机构 * Department of Informatics, University of Sussex(Sussex大学信息学院) Beyond Institute for Theoretical Science (BITS)(理论科学研究所) ETH Zürich(苏黎世联邦理工学院) Principles of Intelligent Behaviour in Biological and Social Systems (PIBBSS)(生物和社会系统智能行为原理研究所) Department of Computer Science, University of Oxford(牛津大学计算机科学系) Araya Inc.(Araya公司) Sussex AI and Sussex Centre for Consciousness Science, University of Sussex(Sussex大学人工智能与意识科学中心) Centre for Complexity Science and Center for Psychedelic Research, Department of Brain Sciences, Imperial College London(复杂科学中心和迷幻研究中心,伦敦帝国理工学院脑科学系) Center for Eudaimonia and Human Flourishing, University of Oxford(幸福与人类繁荣中心,牛津大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出了一种分解复杂world建模的方法,通过转换器框架将世界模型分解为多个模块,从而提高计算效率并支持分布式推理,为AI安全和现实应用提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04309 2026-05-19 cs.LG 70%

Activation Steering with a Feedback Controller

通过反馈控制器激活控制

Dung V. Nguyen, Hieu M. Vu, Nhi Y. Pham, Lei Zhang, Tan M. Nguyen

机构 * Department of Mathematics(数学系) Center for AI Research(人工智能研究中心) National University of Singapore(新加坡国立大学) VinUniversity(文大学) Torilab(Torilab实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本文提出PID激活控制方法,基于控制理论构建激活控制框架,通过P、I、D项实现激活对齐、误差累积和抑制超调,提升大语言模型行为控制的鲁棒性和可靠性。

Comments 10 pages in the main text. ICLR2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01201 2026-05-18 cs.LG cs.CV 70%

Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models

走出洞穴:JAM用于对齐独立训练的视觉和语言模型

Lauren Hyoseo Yoon, Yisong Yue, Been Kim

机构 * Computation and Neural Systems(计算与神经系统) California Institute of Technology(加利福尼亚理工学院) Computation and Mathematical Sciences(计算与数学科学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.LG

AI总结 本文提出JAM方法,通过联合训练模态特定的自编码器,优化视觉和语言模型的对齐,提升细粒度上下文区分能力。

详情

展开后加载摘要…

URL PDF HTML 收藏