arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-07 至 2025-10-07 共收录 82 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 9 篇

2509.10127 2025-10-07 cs.CL cs.AI cs.LG 67%

Population-Aligned Persona Generation for LLM-based Social Simulation

Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Max Xiong, Yuxuan Lei, Tianfu Wang, Kaize Ding, Ziang Xiao, Nicholas Jing Yuan, Xing Xie

机构 * HKUST(香港科技大学) Microsoft Research Asia(微软亚洲研究院) Duke University(杜克大学) Northwestern University(西北大学) Johns Hopkins University(约翰霍普金斯大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18562 2025-10-07 cs.CL cs.AI 66%

From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test

Xunlian Dai, Li Zhou, Benyou Wang, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI

Comments Cultural Analysis, Cultural Alignment, Word Association Test, Large Language Models. Accepted by EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10659 2025-10-07 cs.SI cs.AI cs.CL cs.MA 62%

Network Formation and Dynamics Among Multi-LLMs

Marios Papachristou, Yuan Yuan

机构 * Department of Information Systems, W.P. Carey School of Business, Arizona State University, Tempe, AZ, USA(亚利桑那州立大学信息系统系,W.P. Carey商学院,Tempe分校) Department of Computer Science, Cornell University, Ithaca, NY, USA(康奈尔大学计算机科学系) Graduate School of Management, University of California Davis, Davis, CA, USA(加州大学戴维斯分校管理研究生院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at PNAS Nexus

Journal ref PNAS Nexus 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03368 2025-10-07 cs.CY cs.AI 62%

An Adaptive Responsible AI Governance Framework for Decentralized Organizations

Kiana Jafari Meimandi, Anka Reuel, Gabriela Aranguiz-Dias, Hatim Rahama, Ala-Eddine Ayadi, Xavier Boullier, Jérémy Verdo, Louis Montanie, Mykel Kochenderfer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 57%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04038 2025-10-07 eess.SY cs.SY 50%

Distributed MPC-based Coordination of Traffic Perimeter and Signal Control: A Lexicographic Optimization Approach

Viet Hoang Pham, Hyo-Sung Ahn

专题命中 AI治理与伦理 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他安全 16 篇

2510.03374 2025-10-07 cs.CY cs.AI cs.CL 82%

Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study

Antoun Yaacoub, Zainab Assaghir, Jérôme Da-Rugna

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published in the 36th Central European Conference on Information and Intelligent Systems(CECIIS)at: Varaždin, Croatia. September 17-19/2025. ISSN 1847-2001 (Print). ISSN 1848-2295 (Online)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04919 2025-10-07 cs.CL cs.AI cs.DB 81%

Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment

Davood Rafiei, Morgan Lindsay Heisler, Weiwei Zhang, Mohammadreza Pourreza, Yong Zhang

机构 * University of Alberta(阿尔伯塔大学) Huawei Tech. Canada(华为技术加拿大)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04202 2025-10-07 cs.LG 79%

Spectral Alignment as Predictor of Loss Explosion in Neural Network Training

Haiquan Qiu, You Wu, Yingjie Tan, Yaqing Wang, Quanming Yao

机构 * Tsinghua University(清华大学) Beijing Institute of Mathematical Sciences and Applications(北京数学科学研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04145 2025-10-07 cs.CV cs.CL cs.IR 79%

Automating construction safety inspections using a multi-modal vision-language RAG framework

Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen, Alireza Ahmadian Fard Fini, Daniel Dias-da-Costa

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments 33 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03727 2025-10-07 cs.AI cs.CL cs.CV cs.LG 67%

Bridging the Gap Between Multimodal Foundation Models and World Models

Xuehai He

机构 * Computer Science and Engineering University of California, Santa Cruz(计算机科学与工程大学加州大学圣克ruz分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03845 2025-10-07 cs.AI cs.GT cs.LG stat.ML 62%

The Hidden Game Problem

Gon Buzaglo, Noah Golowich, Elad Hazan

机构 * Princeton University(普林斯顿大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04397 2025-10-07 cs.CR cs.AI cs.SE 57%

MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection

Van Nguyen, Surya Nepal, Xingliang Yuan, Tingmin Wu, Fengchao Chen, Carsten Rudolph

机构 * Monash University(墨尔本大学) CSIRO’s Data61(澳大利亚联邦科学与工业研究组织数据61分部) The University of Melbourne(墨尔本大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04239 2025-10-07 cs.IR cs.AI 57%

Empowering Denoising Sequential Recommendation with Large Language Model Embeddings

Tongzhou Wu, Yuhao Wang, Maolin Wang, Chi Zhang, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Harbin Engineering University(哈尔滨工程大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by CIKM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25694 2025-10-07 cs.SD cs.AI 57%

HNote: Extending YNote with Hexadecimal Encoding for Fine-Tuning LLMs in Music Modeling

Hung-Ying Chu, Shao-Yu Wei, Guan-Wei Chen, Tzu-Wei Hung, ChengYang Tsai, Yu-Cheng Lin

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03645 2025-10-07 cs.AI 57%

Graph Generation Powered with LLMs for Boosting Multivariate Time-Series Representation Learning

Yucheng Wang, Min Wu, Ruibing Jin, Xiaoli Li, Lihua Xie, Zhenghua Chen

机构 * Institute for Infocomm Research, A ∗ STAR, Singapore(信息通信研究所,A*STAR,新加坡) School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(电气电子工程学院,南洋理工大学,新加坡) College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16876 2025-10-07 cs.SE 50%

Revolutionizing Validation and Verification: Explainable Testing Methodologies for Intelligent Automotive Decision-Making Systems

Halit Eris, Stefan Wagner

专题命中 其他安全 :safety(abstract)

Comments Preprint to be published at SE4ADS

Journal ref 2025 IEEE/ACM 1st International Workshop on Software Engineering for Autonomous Driving Systems (SE4ADS), Ottawa, ON, Canada, 2025, pp. 34-37

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20117 2025-10-07 cs.CV 50%

RESCUE: Crowd Evacuation Simulation via Controlling SDM-United Characters

Xiaolin Liu, Tianyi Zhou, Hongbo Kang, Jian Ma, Ziwen Wang, Jing Huang, Wenguo Weng, Yu-Kun Lai, Kun Li

机构 * Tianjin University(天津大学) Tsinghua University(清华大学) Cardiff University(卡迪夫大学)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10666 2025-10-07 astro-ph.SR astro-ph.GA astro-ph.IM 50%

Machine-learning inference of stellar properties using integrated photometric and spectroscopic data

Ilay Kamai, Alex M. Bronstein, Hagai B. Perets

专题命中 其他安全 :alignment(abstract)

Comments Accepted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18569 2025-10-07 cs.CV 50%

VisualChef: Generating Visual Aids in Cooking via Mask Inpainting

Oleh Kuzyk, Zuoyue Li, Marc Pollefeys, Xi Wang

机构 * ETH Zürich(苏黎世联邦理工学院) Microsoft(微软) TU Munich(慕尼黑工业大学) MCML

专题命中 其他安全 :alignment(abstract)

Comments GCPR 2025 (oral presentation; Best Master's Thesis Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11368 2025-10-07 cs.CV 50%

From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation

Jingkun Chen, Haoran Duan, Xiao Zhang, Boyan Gao, Vicente Grau, Jungong Han

机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学) Department of Automation, Tsinghua University(自动化系,清华大学) School of Information Science and Technology, Northwest University(信息科学与技术学院,西北大学)

专题命中 其他安全 :alignment(abstract)

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16357 2025-10-07 cs.CV 50%

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You, Jianbo Yuan, Hongxia Yang, Chenfeng Xu

机构 * Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他安全 :alignment(abstract)

Comments The code is available at https://github.com/bronyayang/Law_of_Vision_Representation_in_MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏