arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-22 至 2025-07-22 共收录 13 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 13 篇

2502.15639 2025-07-22 cs.CL cs.AI cs.LG 82%

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15743 2025-07-22 cs.AI cs.CL cs.HC cs.LG 67%

Towards physician-centered oversight of conversational diagnostic AI

Elahe Vedadi, David Barrett, Natalie Harris, Ellery Wulczyn, Shashir Reddy, Roma Ruparel, Mike Schaekermann, Tim Strother, Ryutaro Tanno, Yash Sharma, Jihyeon Lee, Cían Hughes, Dylan Slack, Anil Palepu, Jan Freyberg, Khaled Saab, Valentin Liévin, Wei-Hung Weng, Tao Tu, Yun Liu, Nenad Tomasev, Kavita Kulkarni, S. Sara Mahdavi, Kelvin Guu, Joëlle Barral, Dale R. Webster, James Manyika, Avinatan Hassidim, Katherine Chou, Yossi Matias, Pushmeet Kohli, Adam Rodman, Vivek Natarajan, Alan Karthikesalingam, David Stutz

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究) Harvard Medical School, Beth Israel Deaconess Medical Center(哈佛医学院,贝塞斯达德acons医学中心)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08622 2025-07-22 cs.AI cs.CL cs.CV 62%

Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models

Donghoon Kim, Minji Bae, Kyuhong Shim, Byonghyo Shim

机构 * Seoul National University(首尔国立大学) Sungkyunkwan University(庆熙大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025 (Official Code: https://github.com/DonghoonKim-1938/VGD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15851 2025-07-22 cs.AI 57%

The Other Mind: How Language Models Exhibit Human Temporal Cognition

Lingyu Li, Yang Yao, Yixu Wang, Chubo Li, Yan Teng, Yingchun Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 12 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10406 2025-07-22 cs.CV cs.AI 57%

RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models

Yijing Lin, Mengqi Huang, Shuhan Zhuang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04510 2025-07-22 cs.SE cs.AI 57%

CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability Detection

Ruijun Feng, Hammond Pearce, Pietro Liguori, Yulei Sui

机构 * School of Computer Science and Engineering, University of New South Wales (UNSW)(新南威尔士大学计算机科学与工程学院) Department of Electrical Engineering and Information Technology, University of Naples Federico II(那不勒斯费德里克二世大学电气工程与信息技术系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Software Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01639 2025-07-22 cs.LG cs.SE 57%

ModelVerification.jl: a Comprehensive Toolbox for Formally Verifying Deep Neural Networks

Tianhao Wei, Hanjiang Hu, Luca Marzari, Kai S. Yun, Peizhi Niu, Xusheng Luo, Changliu Liu

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Verona(威尼斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14260 2025-07-22 physics.optics cs.LG 57%

Automating Experimental Optics with Sample Efficient Machine Learning Methods

Arindam Saha, Baramee Charoensombutamon, Thibault Michel, V. Vijendran, Lachlan Walker, Akira Furusawa, Syed M. Assad, Ben C. Buchler, Ping Koy Lam, Aaron D. Tranter

机构 * Australian National University(澳大利亚国立大学) University of Tokyo(东京大学) Agency for Science, Technology and Research(科技研究局) pi Software(2pi软件)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16824 2025-07-22 cs.CV cs.AI 57%

PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding

Vinh Nguyen

机构 * Uppsala University(乌普萨拉大学) Florida Institute of Technology(佛罗里达理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15147 2025-07-22 cs.LO cs.FL cs.MA 50%

STL-GO: Spatio-Temporal Logic with Graph Operators for Distributed Systems with Multiple Network Topologies

Yiqi Zhao, Xinyi Yu, Bardh Hoxha, Georgios Fainekos, Jyotirmoy V. Deshmukh, Lars Lindemann

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14146 2025-07-22 eess.SP 50%

Estimating Markers of Driving Stress through Multimodal Physiological Monitoring

Kleanthis Avramidis, Emily Zhou, Tiantian Feng, Hossein Hamidi Shishavan, Frederico Marcolino Quintao Severgnini, Danny J. Lohan, Paul Schmalenberg, Ercan M. Dede, Shrikanth Narayanan

专题命中 其他安全 :safety(abstract)

Comments 11 pages, 7 figures, 3 tables. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09442 2025-07-22 cs.CV 50%

Advancing Textual Prompt Learning with Anchored Attributes

Zheng Li, Yibing Song, Ming-Ming Cheng, Xiang Li, Jian Yang

机构 * PCA Lab, VCIP, College of Computer Science, Nankai University(PCA实验室、VCIP、计算机科学学院、南开大学) DAMO Academy, Alibaba Group(达摩院、阿里巴巴集团)

专题命中 其他安全 :alignment(abstract)

Comments ICCV 2025. Code: https://github.com/zhengli97/ATPrompt. Project Page: https://zhengli97.github.io/ATPrompt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18507 2025-07-22 cs.RO 50%

At First Contact: Stiffness Estimation Using Vibrational Information for Prosthetic Grasp Modulation

Anway S. Pimpalkar, Ariel Slepyan, Nitish V. Thakor

机构 * Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰·霍普金斯大学) Department of Electrical and Computer Engineering, Johns Hopkins University(电气与计算机工程系,约翰·霍普金斯大学)

专题命中 其他安全 :safety(abstract)

Comments 5 pages, 7 figures, for IEEE Sensors Letters

详情

展开后加载摘要…

URL PDF HTML 收藏