arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8017 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8017 篇

2409.12097 2024-09-20 cs.CL cs.IR cs.LG cs.SI 76%

Skill matching at scale: freelancer-project alignment for efficient multilingual candidate retrieval

Warren Jouanneau, Marc Palyart, Emma Jouffroy

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08078 2024-08-20 cs.CV cs.AI cs.CL 76%

Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation

Wenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo, Xiang Li, Yixuan Yuan

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024

Journal ref https://aclanthology.org/2024.acl-long.514/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06725 2024-08-14 cs.AI cs.CL cs.CV 76%

Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations

Wei Pang, Ruixue Duan, Jinfu Yang, Ning Li

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments This article has been accepted in CAAI Transactions on Intelligence Technology! Article ID: CIT2_12370, Article DOI: 10.1049/cit2.12370

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06540 2024-08-14 physics.acc-ph cs.AI cs.LG 76%

Dynamic Exclusion of Low-Fidelity Data in Bayesian Optimization for Autonomous Beamline Alignment

Megha R. Narayanan, Thomas W. Morris

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments 12 pages, 6 figure sets

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19594 2024-07-31 cs.CL cs.AI 76%

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01103 2024-06-05 cs.AI cs.HC cs.LG 76%

Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment

Chen Zhang, Qiang He, Zhou Yuan, Elvis S. Liu, Hong Wang, Jian Zhao, Yang Wang

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accept at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15823 2024-01-23 cs.CL cs.AI 76%

Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment

Ahmed ElBakry, Mohamed Gabr, Muhammad ElNokrashy, Badr AlKhamissi

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Proceedings of ArabicNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13627 2023-10-25 cs.CL cs.AI 76%

InstructAlign: High-and-Low Resource Language Alignment via Continual Crosslingual Instruction Tuning

Samuel Cahyawijaya, Holy Lovenia, Tiezheng Yu, Willy Chung, Pascale Fung

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16424 2023-09-29 cs.CL cs.AI cs.SI 76%

Prompt-and-Align: Prompt-Based Social Alignment for Few-Shot Fake News Detection

Jiaying Wu, Shen Li, Ailin Deng, Miao Xiong, Bryan Hooi

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted to CIKM 2023 (Full Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05831 2023-09-13 cs.LG cs.AI 76%

Studying Accuracy of Machine Learning Models Trained on Lab Lifting Data in Solving Real-World Problems Using Wearable Sensors for Workplace Safety

Joseph Bertrand, Nick Griffey, Ming-Lun Lu, Rashmi Jha

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments 7 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16550 2023-05-29 cs.CL cs.AI cs.NE 76%

Soft Alignment Objectives for Robust Adaptation of Language Generation

Michal Štefánik, Marek Kadlčík, Petr Sojka

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Annual Meeting of The ACL 2023: Main conference long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06386 2023-05-12 cs.CV cs.AI cs.HC cs.LG 76%

Text-To-Concept (and Back) via Cross-Model Alignment

Mazda Moayeri, Keivan Rezaei, Maziar Sanjabi, Soheil Feizi

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted to ICML 2023 and CVPR4XAI workshop 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15259 2023-05-10 cs.LG cs.AI cs.CR 76%

On the Impossible Safety of Large AI Models

El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, Sébastien Rouault, John Stephan

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00902 2023-02-06 cs.LG cs.CL cs.CV 76%

Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment

Hao Liu, Wilson Yan, Pieter Abbeel

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

Comments Fixed typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03525 2022-04-08 cs.LG cs.AI 76%

Temporal Alignment for History Representation in Reinforcement Learning

Aleksandr Ermolov, Enver Sangineto, Nicu Sebe

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments ICPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13411 2022-03-28 cs.RO cs.AI cs.LG cs.SY eess.SY 76%

Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using Transformers

Arthur Bucker, Luis Figueredo, Sami Haddadin, Ashish Kapoor, Shuang Ma, Rogerio Bonatti

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14659 2021-03-30 cs.AI cs.LG 76%

Alignment of Language Agents

Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, Geoffrey Irving

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06907 2019-10-16 cs.LG cs.AI math.OC 76%

Techniques for Adversarial Examples Threatening the Safety of Artificial Intelligence Based Systems

Utku Kose

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments International Science and Innovation Congress 2019, pp. 643-655, 13 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.01267 2018-11-06 cs.AI cs.CY cs.HC 76%

Legible Normativity for AI Alignment: The Value of Silly Rules

Dylan Hadfield-Menell, McKane Andrus, Gillian K. Hadfield

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22226 2026-06-23 cs.GT cs.AI cs.IT math.IT 新提交 76%

Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion

量化理论上的AI对齐保证:贝叶斯说服中的接收者效用界

Eric Yachbes, Eva Tardos

机构 * Cornell University(康奈尔大学)

专题命中 其他安全 :alignment(title,comments);分类 cs.AI

AI总结 通过贝叶斯说服模型,研究AI发送者优化错位目标时,人类接收者仍能获得多少有用信息,证明接收者效用比不超过3/2,并给出紧性下界。

Comments 12 pages, EC 2026 Poster and EC 2026 Incentive-Based AI Alignment Workshop Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19586 2025-10-28 cs.CL q-bio.NC 76%

Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment

Hanlin Wu, Xufeng Duan, Zhenguang Cai

机构 * Department of Linguistics and Modern Languages, The Chinese University of Hong Kong(语言学与现代语言系,香港中文大学) Brain and Mind Institute, The Chinese University of Hong Kong(脑与心智研究所,香港中文大学)

专题命中 其他安全 :alignment(title,comments);分类 cs.CL

Comments Hanlin Wu, Xufeng Duan, and Zhenguang Cai. 2025. Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 135-143, Albuquerque, New Mexico, USA. Association for Computational Linguistics. https://aclanthology.org/2025.cmcl-1.18/

Journal ref In Proceedings of CMCL, pages 135-143, ACL (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09544 2026-08-25 cs.CL cs.AI cs.LG 版本更新 75%

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

大语言模型通过一种独特的统一机制生成有害内容

Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov

机构 * Kempner Institute, Harvard University(哈佛大学肯普纳研究所) Princeton University(普林斯顿大学) Harvard University(哈佛大学) Cohere Technion—IIT(以色列理工学院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过权重剪枝揭示大语言模型中有害生成的内部结构,发现有害内容生成依赖于一组通用且与良性能力不同的权重,表明对齐训练重塑了有害表示,解释了领域微调引发的广泛对齐偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22925 2026-07-28 cs.CL cs.AI cs.LG 新提交 75%

Not All LLM Reasoning is Visible in the Chain-of-Thought

并非所有大语言模型的推理都能在思维链中体现

Vatsal Baherwani, Tom Goldstein, Ashwinee Panda

机构 * New York University(纽约大学) University of Maryland(马里兰大学) TogetherAI

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨大语言模型输出令牌是否体现所有推理,发现前沿模型存在利用无关填充令牌提升合成推理任务性能的不可见推理现象,评估多个模型,揭示填充令牌益处因模型和令牌而异,还表明其能服务隐藏目标,且强化学习等方法无法使填充令牌益处在测试时持续。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04885 2026-06-15 cs.CL cs.AI cs.LG 版本更新 75%

CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters

CuMA: 通过人口统计感知的适配器混合使大语言模型与稀疏文化价值观对齐

Ao Sun, Xiaoyu Wang, Zhe Tan, Yu Li, Jiachen Zhu, Yuheng Jia, Shu Su

机构 * Southeast University(东南大学) ByteDance Inc.(字节跳动公司) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),中华人民共和国教育部,中国)

专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出CuMA框架,通过人口统计感知路由将冲突梯度分离到专家子空间,解决密集模型在多文化对齐中的均值崩溃问题,在WorldValuesBench等基准上取得最优性能。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22005 2026-05-26 cs.LG cs.AI cs.CL 75%

Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)

检查你的大语言模型的秘密词典!五行代码揭示你的大语言模型学到了什么(包括它不应该学到的)

Hisashi Miyashita

机构 * Mgnite Inc.(Mgnite公司)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过对lm_head权重矩阵进行奇异值分解(仅需五行PyTorch代码且无需模型推理),直接从模型权重中揭示可解释的语义子空间,并发现模型训练数据组成和策展哲学。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02064 2026-05-19 cs.LG cs.AI cs.CL 75%

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct

对Llama3-8b-Instruct自生成文本识别能力的检查与控制

Christopher Ackerman, Nina Panickssery

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究探讨了LLM是否能识别自身生成的文本,发现Llama3-8b-Instruct模型能够区分自身输出与人类输出,并通过残差流中的特定向量控制其行为和感知,揭示了模型自我归属的认知机制。

Comments 10 pages, 13 figs, 2 tables, accepted as conference paper to ICLR 2025

Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20164 2026-05-12 cs.LG cs.AI cs.CL 75%

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

计划是什么?LLMs中隐式规划的度量及其在押韵生成和问答中的应用

Jim Maar, Denis Paperno, Callum Stuart McDougall, Neel Nanda

机构 * HPI / University of Potsdam(HPI/波茨坦大学) Utrecht University(乌特勒支大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出简单方法评估LLM隐式规划,通过押韵生成和问答案例展示其可扩展性,发现隐式规划在1B参数模型中普遍存在,为AI安全提供新视角。

Comments 41 pages, 34 figures, Accepted at ICLR 2026, Code available at https://github.com/Jim-Maar/implicit-planning-in-llms

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17614 2026-04-21 cs.AI cs.CL cs.LG 75%

Characterizing Model-Native Skills

刻画模型内禀技能

Feiyang Kang, Mahavir Dabas, Myeongseob Ko, Ruoxi Jia

机构 * Virginia Tech(弗吉尼亚理工学院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出模型内禀技能的刻画方法,通过从序列激活中恢复紧凑正交基,实现行为变化轴的自组织,验证了在推理和安全对齐中的有效性,优于人类定义的技能。

Comments We argue that when the goal is to intervene on model behavior, skill characterization should be *model-native*: grounded in the model's own representations rather than imposed through external ontologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03230 2026-03-04 cs.LG cs.AI cs.CL math.OC 75%

DiaBlo: Diagonal Blocks Are Sufficient For Finetuning

DiaBlo: 对角块足以用于微调

Selcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong, Xin Li, Naigang Wang, Penghang Yin, Zi Yang

机构 * University at Albany, SUNY(纽约州立大学阿尔巴尼分校) IBM T. J. Watson Research Center(IBM 汤普逊·杰·沃森研究中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 DiaBlo是一种仅更新模型权重矩阵对角块的参数高效微调方法,通过消除低秩矩阵乘积需求,实现稳定收敛和高效训练。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24940 2026-01-01 cs.AI cs.CL cs.LG 75%

Iterative Deployment Improves Planning Skills in LLMs

迭代部署提升大语言模型的规划能力

Augusto B. Corrêa, Yoav Gelberg, Luckeciano C. Melo, Ilia Shumailov, André G. Pereira, Yarin Gal

机构 * University of Oxford(牛津大学) AI Sequrity Company(AI安全公司) UFRGS(乌拉圭联邦大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过迭代部署大语言模型,利用用户编纂的数据提升规划能力,展现隐含奖励函数的强化学习机制,具有AI安全和训练制度替代的双重意义。

详情

展开后加载摘要…

URL PDF HTML 收藏