arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-23 至 2025-09-23 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 10 篇

2509.17943 2025-09-23 cs.CV cs.LG 79%

Can multimodal representation learning by alignment preserve modality-specific information?

Romain Thoreau, Jessie Levillain, Dawa Derksen

机构 * institutetext(机构文本) CNES(法国国家空间研究中心) INSA-IMT(法国里尔INSA-IMT)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Accepted as a workshop paper at MACLEAN - ECML/PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17074 2025-09-23 cs.CV cs.AI 79%

Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models

Qian Zhang, Lin Zhang, Xing Fang, Mingxin Zhang, Zhiyuan Wei, Ran Song, Wei Zhang

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Submitted to the IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16633 2025-09-23 cs.CV cs.AI cs.CL 76%

When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs

Abhirama Subramanyam Penamakuri, Navlika Singh, Piyush Arora, Anand Mishra

机构 * Indian Institute of Technology Jodhpur(印度理工学院朱诺尔)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted to EMNLP (Main) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16660 2025-09-23 cs.CL 70%

Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation

Zuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md Shad Akhtar

机构 * IIIT Dharwad, India(印度IIIT达尔瓦德大学) IIIT Delhi, India(印度IIIT德里大学) FLaME-NLP Lab, IIIT Delhi(IIIT德里大学FLaME-NLP实验室)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments Accepted to the NeurIPS 2025 Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00742 2025-09-23 cs.CL cs.LG 62%

Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents

Sarah Mercer, Daniel P. Martin, Phil Swatton

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17224 2025-09-23 q-bio.BM cs.LG physics.bio-ph 57%

AI-based Methods for Simulating, Sampling, and Predicting Protein Ensembles

Bowen Jing, Bonnie Berger, Tommi Jaakkola

机构 * CSAIL, Massachusetts Institute of Technology(CSAIL,麻省理工学院) Dept. of Mathematics, Massachusetts Institute of Technology(数学系,麻省理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17167 2025-09-23 cs.CL 57%

SFT-TA: Supervised Fine-Tuned Agents in Multi-Agent LLMs for Automated Inductive Thematic Analysis

Seungjun Yi, Joakim Nguyen, Huimin Xu, Terence Lim, Joseph Skrovan, Mehak Beri, Hitakshi Modi, Andrew Well, Liu Leqi, Mia Markey, Ying Ding

机构 * University of Texas at Austin(德克萨斯大学) Vanderbilt University School of Medicine(范德比大学医学院) McCombs School of Business University of Texas at Austin(德克萨斯大学麦康姆商学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11528 2025-09-23 cs.AI 57%

A Survey of Personalized Large Language Models: Progress and Future Directions

Jiahong Liu, Zexuan Qiu, Zhongyang Li, Quanyu Dai, Wenhao Yu, Jieming Zhu, Minda Hu, Menglin Yang, Tat-Seng Chua, Irwin King

机构 * The Chinese University of Hong Kong(香港中文大学) Huawei Technologies Co., Ltd(华为技术有限公司) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) National University of Singapore(新加坡国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 34 pages, 8 figures, 7 tables, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16956 2025-09-23 cs.CV 50%

VidCLearn: A Continual Learning Approach for Text-to-Video Generation

Luca Zanchetta, Lorenzo Papa, Luca Maiano, Irene Amerini

机构 * Sapienza University of Rome, Italy(罗马大学萨皮恩扎) ESA philab(欧洲航天局实验室)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16946 2025-09-23 physics.optics 50%

Machine learning meets Singular Optics II: Single-pixel Detection of Structured Light

Purnesh Singh Badavath, Vijay Kumar

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏