arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-24 至 2025-09-24 共收录 53 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 10 篇

2507.14944 2025-09-24 cs.HC 67%

LEKIA: Expert-Aligned AI Behavior Design for High-Risk Human-AI Interactions

Boning Zhao, Yutong Hu, Xinnuo Li

专题命中 安全评测 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18221 2025-09-24 cs.AI cs.LG 62%

Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00046 2025-09-24 cs.CV cs.AI 57%

Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity

Zhaoyi Joey Hou, Adriana Kovashka, Xiang Lorraine Li

机构 * Department of Computer Science University of Pittsburgh(计算机科学系宾夕法尼亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments To Appear in EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18568 2025-09-24 cs.LG 57%

Explainable Graph Neural Networks: Understanding Brain Connectivity and Biomarkers in Dementia

Niharika Tewari, Nguyen Linh Dan Le, Mujie Liu, Jing Ren, Ziqi Xu, Tabinda Sarwar, Veeky Baths, Feng Xia

机构 * School of Computing Technologies RMIT University Melbourne VIC Australia Department of Biological Sciences Department of Computer Science \& Information Systems Birla Institute of Technology Institute of Innovation, Science Sustainability Federation University Australia Ballarat VIC Australia RMIT University Birla Institute of Technology Federation University Australia

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18869 2025-09-24 cs.DC 50%

On The Reproducibility Limitations of RAG Systems

Baiqiang Wang, Dongfang Zhao, Nathan R Tallent, Luanzheng Guo

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17537 2025-09-24 cs.CV 50%

SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

Dian Jin, Yanghao Zhou, Jinxing Zhou, Jiaqi Ma, Ruohao Guo, Dan Guo

专题命中 安全评测 :alignment(abstract)

Comments Project page: https://github.com/DianJin-HFUT/SimToken

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他安全 17 篇

2406.06144 2025-09-24 cs.CL cs.AI 81%

Language Models Resist Alignment: Evidence From Data Compression

Jiaming Ji, Kaile Wang, Tianyi Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Juntao Dai, Yunhuai Liu, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Computer Science, Peking University(北京大学计算机科学学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10064 2025-09-24 cs.NE cs.LG q-bio.NC 79%

Dynamical Alignment: A Principle for Adaptive Neural Computation

Xia Chen

机构 * Georg Nemetschek Institute Munich Data Science Institute(慕尼黑数据科学研究所) Technische Universität München(慕尼黑技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments 16 pages, 10 figures;

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19018 2025-09-24 cs.LG 79%

OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment

Teng Xiao, Zuchao Li, Lefei Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18369 2025-09-24 cs.CV cs.AI 79%

Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning

Riad Ahmed Anonto, Sardar Md. Saffat Zabin, M. Saifur Rahman

机构 * Bangladesh University of Engineering and Technology (BUET)(孟加拉工程与技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18439 2025-09-24 cs.CL cs.AI 62%

Developing an AI framework to automatically detect shared decision-making in patient-doctor conversations

Oscar J. Ponce-Ponte, David Toro-Tobon, Luis F. Figueroa, Michael Gionfriddo, Megan Branda, Victor M. Montori, Saturnino Luz, Juan P. Brito

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 53 pages, 1 figure, 4 tables, 5 supplementary figures, 13 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18113 2025-09-24 cs.CL cs.LG 62%

Dynamic Prompt Fusion for Multi-Task and Cross-Domain Adaptation in LLMs

Xin Hu, Yue Kang, Guanzi Yao, Tianze Kang, Mengjie Wang, Heyao Liu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12734 2025-09-24 cs.CL cs.AI 62%

Pandora: A Code-Driven Large Language Model Agent for Unified Reasoning Across Diverse Structured Knowledge

Yongrui Chen, Junhao He, Linbo Fu, Shenyu Zhang, Rihui Jin, Xinbang Dai, Jiaqi Li, Dehai Min, Nan Hu, Yuxin Zhang, Guilin Qi, Yi Huang, Tongtong Wu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments New version is arXiv:2508.17905

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04183 2025-09-24 cs.CL cs.AI 62%

GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding

Ziyin Zhang, Hang Yu, Shijie Li, Peng Di, Jianguo Li, Rui Wang

机构 * Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18576 2025-09-24 cs.RO cs.AI 57%

LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA

Zeyi Kang, Liang He, Yanxin Zhang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学中国) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学巴黎法国)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18109 2025-09-24 cs.LG 57%

Machine Learning-Based Classification of Vessel Types in Straits Using AIS Tracks

Jonatan Katz Nielsen

机构 * Copenhagen Business School(哥本哈根商业学院)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18383 2025-09-24 cs.CL 57%

NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities

Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata, Muhammad Abdul-Mageed

机构 * The University of British Columbia(不列颠哥伦比亚大学) Mohammed VI Polytechnic University(穆莱·阿卜杜勒阿齐兹国王理工学院) Tuwaiq Academy(图瓦伊克学院) Invertible AI(可逆人工智能)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main Conference). Camera-ready version. Data & models: https://github.com/UBC-NLP/nilechat

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13707 2025-09-24 cs.CV cs.AI 57%

EventVL: Understand Event Streams via Multimodal Large Language Model

Pengteng Li, Yunfan Lu, Pinghao Song, Wuyang Li, Huizai Yao, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) KU Leuven(根特大学) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院) Carleton University(卡尔顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03382 2025-09-24 cs.GT 50%

Rationality and Behavior Feedback in a Model of Vehicle-to-Vehicle Communication

Brendan Gould, Philip Brown

专题命中 其他安全 :safety(abstract)

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18937 2025-09-24 cs.RO 50%

Lang2Morph: Language-Driven Morphological Design of Robotic Hands

Yanyuan Qiao, Kieran Gilday, Yutong Xie, Josie Hughes

机构 * CREATE Lab, Swiss Federal Institute of Technology Lausanne (EPFL)(瑞士联邦理工学院洛桑分校CREATE实验室) Computer Vision Department, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(马尔代夫比兹莱浙江大学人工智能大学计算机视觉部门)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18839 2025-09-24 cs.CV 50%

Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography

Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

机构 * Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)

专题命中 其他安全 :alignment(abstract)

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18733 2025-09-24 cs.CV 50%

Knowledge Transfer from Interaction Learning

Yilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu, Shugong Xu

机构 * Shanghai University(上海大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18636 2025-09-24 cs.RO 50%

Number Adaptive Formation Flight Planning via Affine Deformable Guidance in Narrow Environments

Yuan Zhou, Jialiang Hou, Guangtong Xu, Fei Gao

机构 * Institute of Cyber-Systems and Control, College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院系统与控制研究所)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏