arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-12 至 2025-08-12 共收录 29 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 29 篇

2508.08131 2025-08-12 cs.CL cs.AI 81%

Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models

Wenze Xu, Chun Wang, Jiazhen Yu, Sheng Chen, Liang Gao, Weihong Deng

机构 * Mashang Consumer Finance Co., Ltd.(Mashang消费金融有限公司) The University of Sydney(悉尼大学) Macau University of Science and Technology(澳门科学技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments To be presented at ACPR 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06592 2025-08-12 cs.CY cs.AI 81%

Towards Integrated Alignment

Ben Y. Reis, William La Cava

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06895 2025-08-12 cs.CV cs.AI 79%

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21298 2025-08-12 cs.SD cs.AI cs.CL cs.LG cs.MM eess.AS 67%

Exploring Adapter Design Tradeoffs for Low Resource Music Generation

Atharva Mehta, Shivam Chauhan, Monojit Choudhury

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14701 2025-08-12 cs.CL cs.AI cs.HC cs.LG q-bio.NC 67%

COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling

Baihan Lin, Djallel Bouneffouf, Yulia Landa, Rachel Jespersen, Cheryl Corcoran, Guillermo Cecchi

机构 * Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai(人工智能与人类健康系,伊坎医学院 Mount Sinai 分校) Department of Psychiatry, Icahn School of Medicine at Mount Sinai(精神病学系,伊坎医学院 Mount Sinai 分校) Department of Neuroscience, Icahn School of Medicine at Mount Sinai(神经科学系,伊坎医学院 Mount Sinai 分校) Berkman Klein Center for Internet & Society, Harvard University(互联网与社会研究中心,哈佛大学) IBM Research, T.J. Watson Research Center(IBM 研究,T.J. Watson 研究中心) Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center(精神疾病研究、教育与临床中心,James J. Peters VA 医疗中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Translational Psychiatry, in press. This work extends our research series in computational psychiatry (e.g auto annotation in arXiv:2204.05522, topic extraction in arXiv:2204.10189, and diagnosis in arXiv:2210.15603) with the introduction of LLMs to complete the full cycle of interpreting and understanding psychotherapy strategies as a comprehensive analytical framework

Journal ref Transl Psychiatry 15, 166 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07484 2025-08-12 cs.CL cs.AI 62%

ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models

Archchana Sindhujan, Shenbin Qian, Chan Chi Chun Matthew, Constantin Orasan, Diptesh Kanojia

机构 * Institute for People-Centred AI and Centre for Translation Studies, School of Computer Science and Electronic Engineering, University of Surrey(以人为本的人工智能研究所和翻译研究中心,计算机科学与电子工程学院,萨里大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to COLM 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22919 2025-08-12 cs.CL cs.AI 62%

A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations

Qixuan Hu, Xumou Zhang, Jinman Kim, Florence Bourgeois, Adam G. Dunn

机构 * School of Computer Science, Faculty of Engineering, University of Sydney(悉尼大学计算机科学学院、工程学院) Computational Health Informatics Program, Boston Children’s Hospital(波士顿儿童医院计算健康信息学项目) Harvard-MIT Center for Regulatory Science and Department of Pediatrics, Harvard Medical School(哈佛-麻省理工监管科学中心和哈佛医学院儿科部门) Sydney School of Public Health, Faculty of Medicine and Health, University of Sydney(悉尼大学公共卫生学院、医学与健康学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures. Updated to include Table 2, Supplementary Table 1, and an additional baseline random forest model

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21931 2025-08-12 cs.IR cs.AI cs.CL cs.MA 62%

ARAG: Agentic Retrieval Augmented Generation for Personalized Recommendation

Reza Yousefi Maragheh, Pratheek Vadla, Priyank Gupta, Kai Zhao, Aysenur Inan, Kehui Yao, Jianpeng Xu, Praveen Kanumala, Jason Cho, Sushant Kumar

机构 * Walmart Global Tech(沃尔玛全球科技)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07480 2025-08-12 eess.SP cs.AI cs.LG 62%

EEG-Language Pretraining for Highly Label-Efficient Clinical Phenotyping

Sam Gijsen, Kerstin Ritter

机构 * Charité – Universitätsmedizin Berlin, Department of Psychiatry and Psychotherapy, Berlin, Germany(柏林查理医院医学大学精神病与心理治疗系) Hertie Institute for AI in Brain Health, University of Tübingen, Germany(图宾根大学健康人工智能研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01196 2025-08-12 cs.CL cs.AI 62%

$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models

Zian Su, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang

机构 * Purdue University(普渡大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments COLM 2025. The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20758 2025-08-12 stat.AP cs.AI cs.CL 62%

Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth

Seyed Pouyan Mousavi Davoudi, Amin Gholami Davodi, Alireza Amiri-Margavi, Alireza Shafiee Fard, Mahdi Jafari

机构 * Independent Researcher in AI and Statistics(人工智能与统计学独立研究者) Shahrood University of Technology(沙霍罗德大学) University of Pittsburgh(匹兹堡大学) Duquesne University(杜克森大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17797 2025-08-12 cs.AI cs.GT cs.LG cs.MA 62%

Observation Interference in Partially Observable Assistance Games

Scott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart Russell

机构 * Center for Human-Compatible AI, University of California, Berkeley(人类兼容人工智能中心,加州大学伯克利分校) Foundations of Cooperative AI Lab, Carnegie Mellon University(协作人工智能实验室,卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08011 2025-08-12 cs.CL 57%

Progressive Depth Up-scaling via Optimal Transport

Mingzi Cao, Xi Wang, Nikolaos Aletras

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07885 2025-08-12 cs.RO cs.AI cs.CV cs.SY eess.SY 57%

Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning

Shoaib Ahmmad, Zubayer Ahmed Aditto, Md Mehrab Hossain, Noushin Yeasmin, Shorower Hossain

机构 * Department of Mechanical Engineering(机械工程系) Rajshahi University of Engineering and Technology(拉贾沙希工程与技术大学) Department of Industrial and Production Engineering(工业与生产工程系) Shahjalal University of Science and Technology(沙赫jalal科学与技术大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Department of Urban and Regional Planning(城市与区域规划系) Department of Computer Science Engineering(计算机科学与工程系) United International University(联合国际大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09969 2025-08-12 cs.HC cs.AI 57%

Steering AI-Driven Personalization of Scientific Text for General Audiences

Taewook Kim, Dhruv Agarwal, Jordan Ackerman, Manaswi Saha

机构 * Northwestern University(西北大学) Cornell University(康奈尔大学) Center for Advanced AI, Accenture(Accenture高级人工智能中心) Accenture Labs(Accenture实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 28 pages, 7 figures, 1 table. Accepted to PACM HCI (CSCW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06765 2025-08-12 cs.LG 57%

Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning

Xingke Yang, Liang Li, Sicong Li, Liwei Guan, Hao Wang, Xiaoqi Qi, Jiang Liu, Xin Fu, Miao Pan

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07989 2025-08-12 cs.CV cs.HC 50%

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 其他安全 :safety(abstract)

Comments 9 pages, 3 figures, 2 tables. Accepted at CV4A11y, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07880 2025-08-12 cs.MA 50%

Multi-agent systems for chemical engineering: A review and perspective

Sophia Rupprecht, Qinghe Gao, Tanuj Karia, Artur M. Schweidtmann

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07804 2025-08-12 cs.CV 50%

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

Bao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan, Xiangyu Zhu, Zhen Lei

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学院大学) CAIR, HKISI, CAS(中国科学院自动化研究所) Beihang University(北京航空航天大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06891 2025-08-12 eess.IV cs.CV 50%

Fusion-Based Brain Tumor Classification Using Deep Learning and Explainable AI, and Rule-Based Reasoning

Melika Filvantorkaman, Mohsen Piri, Maral Filvan Torkaman, Ashkan Zabihi, Hamidreza Moradi

专题命中 其他安全 :alignment(abstract)

Comments 37 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06845 2025-08-12 cs.CV cs.CE eess.IV 50%

Hybrid Machine Learning Framework for Predicting Geometric Deviations from 3D Surface Metrology

Hamidreza Samadi, Md Manjurul Ahsan, Shivakumar Raman

机构 * Industrial and Systems Engineering University of Oklahoma(工业与系统工程大学俄克拉荷马州立大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10574 2025-08-12 cs.CV cs.MM cs.SD eess.AS 50%

DanceChat: Large Language Model-Guided Music-to-Dance Generation

Qing Wang, Xiaohang Yang, Yilan Dong, Naveen Raj Govindaraj, Gregory Slabaugh, Shanxin Yuan

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22824 2025-08-12 math.GM 50%

Constrained Hamiltonian Systems on Observation-Induced Fiber Bundles: Theory of Symmetry and Integrability

Dongzhe Zheng

专题命中 其他安全 :safety(abstract)

Comments This paper establishes the complete mathematical theory underlying the practical framework presented in "Learning Dynamics under Environmental Constraints via Measurement-Induced Bundle Structures" (ICML 2025 (Forty-Second International Conference on Machine Learning) Spotlight, top 2.5% of ~12,000 submissions)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17306 2025-08-12 cs.HC 50%

Exploring the Temporal Dynamics of Facial Mimicry in Emotion Processing Using Action Units

Meisam Jamshidi Seikavandi, Jostein Fimland, Maria Jung Barrett, Paolo Burelli

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06926 2025-08-12 cs.SE 50%

Integrating Rules and Semantics for LLM-Based C-to-Rust Translation

Feng Luo, Kexing Ji, Cuiyun Gao, Shuzheng Gao, Jia Feng, Kui Liu, Xin Xia, Michael R. Lyu

专题命中 其他安全 :safety(abstract)

Comments Accepted in ICSME 25 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06826 2025-08-12 cs.HC 50%

AdjustAR: AI-Driven In-Situ Adjustment of Site-Specific Augmented Reality Content

Nels Numan, Jessica Van Brummelen, Ziwen Lu, Anthony Steed

专题命中 其他安全 :alignment(abstract)

Comments 4 pages, 1 figure, ACM UIST 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06525 2025-08-12 cs.CV 50%

Large Language Models Facilitate Vision Reflection in Image Classification

Guoyuan An, JaeYoon Kim, SungEui Yoon

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03788 2025-08-12 cs.HC 50%

Frontend Diffusion: Empowering Self-Representation of Junior Researchers and Designers Through Multi-agent System

Zijian Ding, Qinshi Zhang, Mohan Chi, Ziyi Wang

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16102 2025-08-12 eess.AS 50%

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Yifan Yang, Shujie Liu, Jinyu Li, Hui Wang, Lingwei Meng, Haiyang Sun, Yuzhe Liang, Ziyang Ma, Yuxuan Hu, Rui Zhao, Jianwei Yu, Yan Lu, Xie Chen

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏