arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-08 至 2025-08-08 共收录 15 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 15 篇

2508.05616 2025-08-08 cs.LG cs.AI cs.NE cs.RO 62%

TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven Evolution

Zhikai Zhao, Chuanbo Hua, Federico Berto, Kanghoon Lee, Zihan Ma, Jiachen Li, Jinkyoo Park

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2505.04480

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05399 2025-08-08 cs.CV cs.AI cs.LG 62%

UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

Wonjun Kang, Byeongkeun Ahn, Minjae Lee, Kevin Galim, Seunghyuk Oh, Hyung Il Koo, Nam Ik Cho

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Code is available at https://github.com/furiosa-ai/uncage

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05019 2025-08-08 cs.CV cs.AI cs.LG 62%

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes

Sadia Kamal, Tim Oates, Joy Wan

机构 * Department of Computer Science, University of Maryland, Baltimore County(计算机科学系,马里兰大学巴尔的摩县分校) Department of Dermatology, Johns Hopkins University School of Medicine(皮肤科系,约翰霍普金斯大学医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to IJCAI 2025 Workshops. arXiv admin note: substantial text overlap with arXiv:2506.10328

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14964 2025-08-08 cs.CL cs.LG 62%

Efficient Knowledge Injection in LLMs via Self-Distillation

Kalle Kujanpää, Pekka Marttinen, Harri Valpola, Alexander Ilin

机构 * Aalto University(阿alto大学) Finnish Center for Artificial Intelligence (FCAI)(芬兰人工智能中心) System 2 AI

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05427 2025-08-08 cs.AI 57%

Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation

Kartar Kumar Lohana Tharwani, Rajesh Kumar, Sumita, Numan Ahmed, Yong Tang

机构 * International Research Center for Complexity Sciences, Hangzhou International Innovation Institute, Beihang University(国际复杂科学研究中心,杭州创新研究院,北京航空航天大学) School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科学与技术大学) Yangtze Delta Region Institute (Huzhou), University of Electronic Science and Technology of China(长江三角洲地区研究所(湖州),电子科学与技术大学) Government Boys Higher Secondary School, Bukera Sharif, Tando Allahyar, Affiliated with BISE Hyderabad, Sindh, Pakistan(政府男生高级中学,布克里·沙里夫,塔ndo阿勒·耶尔,隶属于Hyderabad BISE,Sindh,巴基斯坦) School of Environment and Architecture, University of Shanghai for Science and Technology(环境与建筑学院,上海科学技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05388 2025-08-08 cs.AI 57%

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal

Silvia García-Méndez, Francisco de Arriba-Pérez, Fátima Leal, Bruno Veloso, Benedita Malheiro, Juan Carlos Burguillo-Rial

机构 * Information Technologies Group, atlanTTic, University of Vigo, Spain(信息科技组,atlanTTic,维戈大学,西班牙) REMIT, Universidade Portucalense, Portugal(REMIT,葡萄牙普拉亚恩斯大学) Faculty of Economics, University of Porto, Portugal(经济学院,波尔图大学,葡萄牙) INESC TEC, Porto, Portugal(INESC TEC,波尔图,葡萄牙) ISEP, Porto, Portugal(ISEP,波尔图,葡萄牙)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05232 2025-08-08 cs.LG 57%

Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs

Feifan Xia, Mingyang Liao, Yuyang Fang, Defang Li, Yantong Xie, Weikang Li, Yang Li, Deguo Xia, Jizhou Huang

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05116 2025-08-08 cs.AI cs.MA 57%

Beyond Automation: Socratic AI, Epistemic Agency, and the Implications of the Emergence of Orchestrated Multi-Agent Learning Architectures

Peer-Benedikt Degen, Igor Asanov

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10479 2025-08-08 cs.AI 57%

DeclareAligner: A Leap Towards Efficient Optimal Alignments for Declarative Process Model Conformance Checking

Jacobo Casas-Ramos, Manuel Lama, Manuel Mucientes

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08800 2025-08-08 cs.CL 57%

Data Processing for the OpenGPT-X Model Family

Nicolo' Brandizzi, Hammam Abdelwahab, Anirban Bhowmick, Lennard Helmer, Benny Jörg Stein, Pavel Denisov, Qasid Saleem, Michael Fromm, Mehdi Ali, Richard Rutmann, Farzad Naderi, Mohamad Saif Agy, Alexander Schwirjow, Fabian Küch, Luzian Hahn, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Dennis Wegener, Nicolas Flores-Herr, Joachim Köhler, Johannes Leveling

机构 * Fraunhofer IAIS(弗劳恩霍夫人工智能研究所) Fraunhofer IIS(弗劳恩霍夫信息处理研究所) DFKI(德意志国防科研机构)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04987 2025-08-08 cs.CV 50%

Unified modality separation: A vision-language framework for unsupervised domain adaptation

Xinyao Li, Jingjing Li, Zhekai Du, Lei Zhu, Heng Tao Shen

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Tongji University(同济大学)

专题命中 其他安全 :alignment(abstract)

Comments Accepted to TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04749 2025-08-08 q-bio.NC 50%

Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia

Yifan Wang, Jingyuan Sun, Jichen Zheng, Yunhao Zhang, Chunyu Ye, Jixing Li, Chengqing Zong, Shaonan Wang

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04585 2025-08-08 eess.AS 50%

UniTalker: Conversational Speech-Visual Synthesis

Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, Haizhou Li

专题命中 其他安全 :alignment(abstract)

Comments 15 pages, 8 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17326 2025-08-08 cs.IR cs.SD eess.AS 50%

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering

Zackary Rackauckas, Julia Hirschberg

机构 * Columbia University(哥伦比亚大学)

专题命中 其他安全 :alignment(abstract)

Comments Accepted to ACL 2025 Workshop MAGMaR

Journal ref Proceedings of the 1st Workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR 2025), pp. 40-46, Vienna, Austria, August 2025. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17963 2025-08-08 cs.CV 50%

M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation

Xiaowei Chi, Junbo Qi, Rongyu Zhang, Shanghang Zhang, Qifeng Liu, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Waseda University(早稻田大学) Peking University(北京大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏