arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8017 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8017 篇

2602.17271 2026-02-20 cs.IT cs.AI math.IT 74%

Federated Latent Space Alignment for Multi-user Semantic Communications

联邦潜在空间对齐用于多用户语义通信

Giuseppe Di Poce, Mario Edoardo Pandolfo, Emilio Calvanese Strinati, Paolo Di Lorenzo

机构 * DIAG Department, Sapienza University of Rome(罗马大学DIAG系) Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT)(国家跨大学电信研究会) CEA Leti, University Grenoble Alpes(CEA Leti,格勒诺布尔大学) DIET Department, Sapienza University of Rome(罗马大学DIET系)

专题命中 其他安全 :alignment(title);分类 cs.AI

AI总结 本文提出联邦潜在空间对齐方法,用于多用户语义通信中的潜在空间对齐,通过共享预均衡器和本地均衡器实现任务导向的高效通信。

Journal ref In 2025 IEEE 26th International Workshop on Signal Processing and Artificial Intelligence for Wireless Communications (SPAWC) (pp. 1-5). IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12878 2026-02-16 cs.CY 74%

Understanding Cultural Alignment in Multilingual LLMs via Natural Debate Statements

通过自然辩论陈述理解多语言大语言模型中的文化契合

Vlad-Andrei Negru, Camelia Lemnaru, Mihai Surdeanu, Rodica Potolea

专题命中 其他安全 :alignment(title);分类 cs.CY

AI总结 本文通过分析自然辩论陈述,揭示了多语言大语言模型在不同文化背景下的价值观差异及其对用户社会文化背景适应能力的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05325 2026-02-16 cs.LG 74%

Quasiparticle Interference Kernel Extraction with Variational Autoencoders via Latent Alignment

通过潜在对齐的变分自编码器提取准粒子干涉核

Yingshuai Ji, Haomin Zhuang, Matthew Toole, James McKenzie, Xiaolong Liu, Xiangliang Zhang

机构 * Department of Computer Science \& \ University of Notre Dame South Bend, IN, USA Department of Physics \& Astronomy University of Notre Dame South Bend, IN, USA

专题命中 其他安全 :alignment(title);分类 cs.LG

AI总结 本文提出基于变分自编码器和潜在对齐的AI框架,实现复杂散射条件下准粒子干涉核的高效提取。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11328 2026-02-13 cs.CL 74%

Evaluating Alignment of Behavioral Dispositions in LLMs

评估大语言模型中行为倾向的对齐情况

Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman, Natalie Harris, Shashir Reddy, Romina Stella, Ariel Goldstein, Marian Croak, Yossi Matias, Amir Feder

机构 * Google(谷歌)

专题命中 其他安全 :alignment(title);分类 cs.CL

AI总结 本研究通过情境判断测试评估大语言模型行为倾向与人类倾向的对齐情况,发现模型在不同共识场景下表现出过度自信、偏离共识及跨模型特质模式,揭示了LLMs声明价值与实际行为之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10138 2026-01-29 cs.LG 74%

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

在具有线性函数逼近的约束马尔可夫决策过程中的episode级安全强化学习的可证明效率

Toshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai, Kazumi Kasaura, Kenta Hoshino, Yohei Hosoe, Yutaka Matsuo

专题命中 其他安全 :safety(title);分类 cs.LG

AI总结 本文提出了一种在约束马尔可夫决策过程中的强化学习算法,实现episode级安全性和$\tilde{\mathcal{O}}(\sqrt{K})$的遗憾,同时具有多项式计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08139 2026-01-14 cs.CV cs.AI 74%

Subspace Alignment for Vision-Language Model Test-time Adaptation

子空间对齐用于视觉-语言模型测试时适应

Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 其他安全 :alignment(title);分类 cs.AI

AI总结 SubTTA通过子空间对齐提升视觉-语言模型测试时适应性能,有效解决模态差距和视觉噪声问题。

Comments 17 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20396 2026-01-08 cs.LG 74%

BiListing: Modality Alignment for Listings

BiListing: 列表模态对齐

Guillaume Guy, Mihajlo Grbovic, Chun How Tan, Han Zhao

机构 * Airbnb

专题命中 其他安全 :alignment(title);分类 cs.LG

AI总结 BiListing通过结合大型语言模型和预训练语言-图像模型,实现列表文本与照片的模态对齐,提升搜索效率并解决冷启动问题。

Journal ref Proceedings of the 34th ACM International Conference on Information and Knowledge Management, CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24826 2026-01-01 cs.CV cs.AI 74%

Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free Control

用于多物体3D场景的2D系统中视频与语言对齐的多信息无导数控制

Jason Armitage, Rico Sennnrich

机构 * University of Zurich(苏黎世大学)

专题命中 其他安全 :alignment(title);分类 cs.AI

AI总结 本文提出了一种无导数优化方法,用于在多物体3D场景中实现视频与语言的对齐,通过在线适应物体遮挡和区分特征来提升跨模态任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05024 2025-10-29 cs.LG 74%

Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment

Nevan Wichers, Aram Ebtekar, Ariana Azarbal, Victor Gillioz, Christine Ye, Emil Ryd, Neil Rathi, Henry Sleight, Alex Mallen, Fabien Roger, Samuel Marks

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments v2 Updates references. v3 Updates references; Adds IFEval results; Improves appendix readability; Adds author contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01454 2025-10-03 cs.CV cs.LG 74%

Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories

Nilay Naharas, Dang Nguyen, Nesihan Bulut, Mohammadhossein Bateni, Vahab Mirrokni, Baharan Mirzasoleiman

机构 * Department of Computer Science, University of California Los Angeles(加州大学洛杉矶分校计算机科学系) Google Research(谷歌研究)

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments 30 pages, 10 figures, 5 tables, link: https://bigml-cs-ucla.github.io/XMAS-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01428 2025-10-03 q-bio.QM cs.AI 74%

BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning

Ching-Huei Tsou, Michal Ozery-Flato, Ella Barkan, Diwakar Mahajan, Ben Shapira

机构 * IBM T.J. Watson Research Center(IBM T.J. Watson研究所以) IBM Research(IBM研究所以)

专题命中 其他安全 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21358 2025-09-29 cs.CV cs.AI 74%

MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification

Jason Jordan, Mohammadreza Akbari Lor, Peter Koulen, Mei-Ling Shyu, Shu-Ching Chen

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments Word count: 5157, Table count: 2, Figure count: 5

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21247 2025-09-26 cs.CV cs.AI 74%

Learning to Look: Cognitive Attention Alignment with Vision-Language Models

Ryan L. Yang, Dipkamal Bhusal, Nidhi Rastogi

机构 * Brown University(布朗大学) Rochester Institute of Technology(罗切斯特理工大学)

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments 7 pages, neurips workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08702 2025-09-11 stat.ME cs.CY stat.AP 74%

Generative AI as a Safety Net for Survey Question Refinement

Erica Ann Metheney, Lauren Yehle

专题命中 其他安全 :safety(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04876 2025-09-08 cs.AI 74%

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration

Jusheng Zhang, Yijia Fan, Kaitong Cai, Xiaofei Sun, Keze Wang

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments Accepted at EMNLP 2025 (Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19574 2025-08-28 cs.CV cs.AI 74%

Multimodal Prototype Alignment for Semi-supervised Pathology Image Segmentation

Mingxi Fu, Fanglei Fu, Xitong Ling, Huaitian Yuan, Tian Guan, Yonghong He, Lianghui Zhu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 其他安全 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06992 2025-08-22 cs.CV cs.AI 74%

MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation

Qilong Xing, Zikai Song, Youjia Zhang, Na Feng, Junqing Yu, Wei Yang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16680 2025-07-25 cs.LG cs.IT cs.NI math.IT 74%

Latent Space Alignment for AI-Native MIMO Semantic Communications

Mario Edoardo Pandolfo, Simone Fiorellino, Emilio Calvanese Strinati, Paolo Di Lorenzo

机构 * DIAG Department, Sapienza University of Rome(罗马萨皮恩扎大学DIAG系) Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT)(国家跨大学电信合作组织(CNIT)) CEA Leti, University Grenoble Alpes(CEA Leti,格勒诺布尔阿尔卑斯大学) DIET Department, Sapienza University of Rome(罗马萨皮恩扎大学DIET系)

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments Proc. of IEEE IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16450 2025-07-24 cs.LG eess.SP 74%

RIS-aided Latent Space Alignment for Semantic Channel Equalization

Tomás Hüttebräucker, Mario Edoardo Pandolfo, Simone Fiorellino, Emilio Calvanese Strinati, Paolo Di Lorenzo

机构 * CEA Leti, University Grenoble Alpes(CEA Leti,格勒诺布尔大学) DIAG Department, Sapienza University of Rome(罗马萨皮恩扎大学DIAG系) Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT)(全国大学电信联合体) DIET Department, Sapienza University of Rome(罗马萨皮恩扎大学DIET系)

专题命中 其他安全 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13216 2025-06-17 cs.CL 74%

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law

Qiming Ge, Shuhao Xing, Songyang Gao, Yunhua Zhou, Yicheng Zou, Songyang Zhang, Zhi Chen, Hang Yan, Qi Zhang, Qipeng Guo, Kai Chen

机构 * Shanghai AI Laboratory(上海人工智能实验室) College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments 9 pages, 9 figures, ACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00942 2025-06-17 cs.CL 74%

Experiential Semantic Information and Brain Alignment: Are Multimodal Models Better than Language Models?

Anna Bavaresco, Raquel Fernández

机构 * Institute for Logic, Language and Computation University of Amsterdam(逻辑、语言与计算研究所 阿姆斯特丹大学)

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Accepted to CoNLL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03523 2025-06-05 cs.CL 74%

TokAlign: Efficient Vocabulary Adaptation via Token Alignment

Chong Li, Jiajun Zhang, Chengqing Zong

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS, Beijing, China(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京)

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments ACL 2025, our codes and models are available at https://github.com/ZNLP/TokAlign

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00168 2025-05-20 cs.CL cs.SD eess.AS 74%

SSR: Alignment-Aware Modality Connector for Speech Language Models

Weiting Tan, Hirofumi Inaguma, Ning Dong, Paden Tomasello, Xutai Ma

机构 * Johns Hopkins University(约翰霍普金斯大学) Meta AI Research(Meta AI研究)

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments IWSLT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02124 2025-05-06 cs.LG 74%

GRAIL: Graph Edit Distance and Node Alignment Using LLM-Generated Code

Samidha Verma, Arushi Goyal, Ananya Mathur, Ankit Anand, Sayan Ranu

专题命中 其他安全 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13840 2025-04-15 cs.IR cs.AI 74%

Multi-view Intent Learning and Alignment with Large Language Models for Session-based Recommendation

Shutong Qiao, Wei Zhou, Junhao Wen, Chen Gao, Qun Luo, Peixuan Chen, Yong Li

专题命中 其他安全 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09353 2025-04-02 cs.LG cs.CV 74%

Enhancing Domain Adaptation through Prompt Gradient Alignment

Hoang Phan, Lam Tran, Quyen Tran, Trung Le

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18603 2025-03-27 cs.CL 74%

LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment

Jong Myoung Kim, Young-Jun Lee, Ho-Jin Choi, Sangkeun Jung

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments now preparing

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15003 2025-03-20 cs.CL 74%

LLM Alignment for the Arabs: A Homogenous Culture or Diverse Ones?

Amr Keleg

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Accepted to the C3NLP workshop (Co-located with NAACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09774 2025-03-14 cs.CL 74%

Efficient Multi-Task Inferencing: Model Merging with Gromov-Wasserstein Feature Alignment

Luyang Fang, Ehsan Latif, Haoran Lu, Yifan Zhou, Ping Ma, Xiaoming Zhai

专题命中 其他安全 :alignment(title);分类 cs.CL

Comments Submitted to AIED2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05087 2025-03-10 cs.LG 74%

Partial Distribution Alignment via Adaptive Optimal Transport

Pei Yang, Qi Tan, Guihua Wen

专题命中 其他安全 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏