arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2508.19099 2025-08-27 cs.CL 57%

Beyond the Black Box: Integrating Lexical and Semantic Methods in Quantitative Discourse Analysis with BERTopic

Thomas Compton

机构 * University of York(约克大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 5 pages conference paper, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19089 2025-08-27 cs.CL 57%

It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs

Yue Li, Zhixue Zhao, Carolina Scarton

机构 * Department of Computer Science, University of Sheffield, UK(计算机科学系,谢菲尔德大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18954 2025-08-27 cs.LG 57%

On the Generalisation of Koopman Representations for Chaotic System Control

Kyriakos Hjikakou, Juan Diego Cardenas Cartagena, Matthia Sabatelli

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18763 2025-08-27 cs.AI 57%

Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units

Chao Hao, Zezheng Wang, Yanhua Huang, Ruiwen Xu, Wenzhe Niu, Xin Liu, Zitong Yu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18381 2025-08-27 cs.CL 57%

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models

Yuchun Fan, Yilin Wang, Yongyu Mu, Lei Huang, Bei Li, Xiaocheng Feng, Tong Xiao, Jingbo Zhu

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17751 2025-08-26 cs.LG 57%

Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning

Alessio Arcudi, Davide Sartor, Alberto Sinigaglia, Vincent François-Lavet, Gian Antonio Susto

机构 * Human Inspired Technology Research Center, Università di Padova, Padova, PD 35121 IT(人类启发技术研究中心,帕多瓦大学,帕多瓦,PD 35121 IT) Vrije Universiteit Amsterdam, Amsterdam, Netherlands(阿姆斯特丹自由大学,阿姆斯特丹,荷兰)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17638 2025-08-26 cs.CV cs.CL 57%

Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning

Xinyu Wei, Guoli Yang, Jialu Zhou, Mingyue Yang, Leqian Li, Kedi Zhang, Chunping Qiu

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17164 2025-08-26 cs.CL 57%

The Impact of Annotator Personas on LLM Behavior Across the Perspectivism Spectrum

Olufunke O. Sarumi, Charles Welch, Daniel Braun, Jörg Schlötterer

机构 * University of Marburg(马尔堡大学) McMaster University(麦马斯特大学) University of Mannheim(曼海姆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at ICNLSP 2025, Odense, Denmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17104 2025-08-26 cs.AI 57%

Rethinking How AI Embeds and Adapts to Human Values: Challenges and Opportunities

Sz-Ting Tzeng, Frank Dignum

机构 * Ume University(乌梅大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 7 pages, accepted at VALE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16962 2025-08-26 cs.RO cs.AI 57%

LLM-based Human-like Traffic Simulation for Self-driving Tests

Wendi Li, Hao Wu, Han Gao, Bing Mao, Fengyuan Xu, Sheng Zhong

机构 * National Key Lab for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09497 2025-08-26 cs.CL 57%

GoalfyMax: A Protocol-Driven Multi-Agent System for Intelligent Experience Entities

Siyi Wu, Zeyu Wang, Xinyuan Song, Zhengpeng Zhou, Lifan Sun, Tianyu Shi

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments The author information is incorrect, some contributors are not included, and the submission has not been approved by all authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21397 2025-08-26 cs.CL 57%

DecisionFlow: Advancing Large Language Model as Principled Decision Maker

Xiusi Chen, Shanyong Wang, Cheng Qian, Hongru Wang, Peixuan Han, Heng Ji

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Findings; 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10624 2025-08-26 cs.CL 57%

SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition

Zechen Li, Shohreh Deldari, Linyao Chen, Hao Xue, Flora D. Salim

机构 * University of New South Wales(新南威尔士大学) University of Tokyo(东京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16776 2025-08-26 cs.LG 57%

Latent Graph Learning in Generative Models of Neural Signals

Nathan X. Kodama, Kenneth A. Loparo

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16696 2025-08-26 cs.GR cs.AI 57%

DecoMind: A Generative AI System for Personalized Interior Design Layouts

Reema Alshehri, Rawan Alotaibi, Leen Almasri, Rawan Altaweel

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments ~7 pages; ~32 figures; compiled with pdfLaTeX. Primary category: cs.CV. (Secondary: cs.AI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16352 2025-08-25 cs.AI eess.SP 57%

Causal Beam Selection for Reliable Initial Access in AI-driven Beam Management

Nasir Khan, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Sinem Coleri

机构 * Department of Electrical and Electronics Engineering, Koç University(电子与电气工程系,科克大学) CEMSE Division, King Abdullah University of Science and Technology(KAUST能源科学与工程分校,国王 Abdullah 科学技术大学) School of Electronics and Computer Science, University of Southampton(电子与计算机科学学院,南安普顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16244 2025-08-25 cs.LG 57%

When Simpler Wins: Facebooks Prophet vs LSTM for Air Pollution Forecasting in Data-Constrained Northern Nigeria

Habeeb Balogun, Yahaya Zakari

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06360 2025-08-25 cs.CL 57%

Cyberbullying Detection via Aggression-Enhanced Prompting

Aisha Saeid, Anu Sabu, Girish A. Koushik, Ferrante Neri, Diptesh Kanojia

机构 * NICE Research Group & Institute for People-Centred AI, School of Computer Science & Electronic Engineering, University of Surrey, UK(NICE研究组及以人为本的人工智能研究所、计算机科学与电子工程学院、萨里大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Accepted to RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15262 2025-08-22 cs.IR cs.AI 57%

M-$LLM^3$REC: A Motivation-Aware User-Item Interaction Framework for Enhancing Recommendation Accuracy with LLMs

Lining Chen, Qingwen Zeng, Huaming Chen

机构 * The University of Sydney(悉尼大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 10pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15346 2025-08-22 q-bio.QM cs.LG 57%

Drug-Target Interaction/Affinity Prediction: Deep Learning Models and Advances Review

Ali Vefghi, Zahed Rahmati, Mohammad Akbari

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 64 pages, 7 figures, 10 tables

Journal ref Journal of Computers in Biology and Medicine Volume 196, Part A, September 2025, 110438

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14618 2025-08-21 cs.LG 57%

A Fuzzy-Enhanced Explainable AI Framework for Flight Continuous Descent Operations Classification

Amin Noroozi, Sandaruwan K. Sethunge, Elham Norouzi, Phat T. Phan, Kavinda U. Waduge, Md. Arafatur Rahman

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14214 2025-08-21 cs.AI 57%

Large Language Models are Highly Aligned with Human Ratings of Emotional Stimuli

Mattson Ogg, Chace Ashcraft, Ritwik Bose, Raphael Norman-Tenazas, Michael Wolmetz

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06723 2025-08-20 cs.HC cs.AI 57%

Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for Psychotherapy

Xin Sun, Jan de Wit, Zhuying Li, Jiahuan Pei, Abdallah El Ali, Jos A. Bosch

机构 * University of Amsterdam(阿姆斯特丹大学) National Institute of Informatics (NII)(日本信息处理研究所) Tilburg University(蒂尔堡大学) Southeast University(东南大学) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) Utrecht University(乌得勒支大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12872 2025-08-19 cs.DB cs.CY 57%

Evaluating the Quality of Open Building Datasets for Mapping Urban Inequality: A Comparative Analysis Across 5 Cities

Franz Okyere, Meng Lu, Ansgar Brunn

专题命中 其他安全 :alignment(abstract);分类 cs.CY

Comments 25 pages, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10935 2025-08-19 cs.CV cs.LG cs.RO 57%

HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model

Qi Liu, Yabei Li, Hongsong Wang, Lei He

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12488 2025-08-19 cs.HC cs.AI 57%

Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing Process

Mohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry, Joseph Jay Williams, Anastasia Kuzminykh

机构 * University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Journal ref PACMHCI (CSCW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11280 2025-08-19 cs.CL 57%

High-Dimensional Interlingual Representations of Large Language Models

Bryan Wilie, Samuel Cahyawijaya, Junxian He, Pascale Fung

机构 * Hong Kong University of Science and Technology(香港理工大学) Cohere

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03885 2025-08-19 cs.LG 57%

Seldonian Reinforcement Learning for Ad Hoc Teamwork

Edoardo Zorzi, Alberto Castellini, Leonidas Bakopoulos, Georgios Chalkiadakis, Alessandro Farinelli

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Presented at the 2nd Reinforcement Learning Conference (RLC2025), Edmonton, Canada. To be published in the Proceedings of the Reinforcement Learning Journal 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11889 2025-08-19 cs.CL 57%

In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning

Hui Ma, Bo Zhang, Jinpeng Hu, Zenglin Shi

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11536 2025-08-18 cs.CL 57%

Language models align with brain regions that represent concepts across modalities

Maria Ryskina, Greta Tuckute, Alexander Fung, Ashley Malkin, Evelina Fedorenko

机构 * Vector Institute for AI(向量人工智能研究所) MIT(麻省理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted to COLM 2025. Code and data can be found at https://github.com/ryskina/concepts-brain-llms

详情

展开后加载摘要…

URL PDF HTML 收藏