arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-10 至 2025-11-10 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8 篇

2410.02615 2025-11-10 cs.LG 79%

ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models

Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen, Hoang-Bao Le, Tai Nguyen, Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, James Zou, Daniel Sonntag, Mathias Niepert

机构 * German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心) Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校) University of Stuttgart(斯图加特大学) University Medical Center Gottingen(哥廷根大学医学中心) Max Planck Institute for Multidisciplinary Sciences(马克斯·普朗克多学科科学研究所) ARC Centre of Excellence for the Mathematical Analysis of Cellular Systems(细胞系统数学分析卓越中心) School of Mathematical Sciences, Queensland University of Technology(昆士兰科技大学数学科学学院) University of Oldenburg(奥尔登堡大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California San Diego(加州大学圣地亚哥分校) MBZUAI(马克斯·普朗克人工智能研究所) ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05718 2025-11-10 cs.HC 71%

Do Vision-Language Models See Visualizations Like Humans? Alignment in Chart Categorization

Péter Ferenc Gyarmati, Manfred Klaffenböck, Laura Koesten, Torsten Möller

专题命中 其他安全 :alignment(title)

Comments 2 pages, 2 figures. Accepted submission to the poster track of IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04995 2025-11-10 cs.HC cs.AI cs.CL 62%

Enhancing Public Speaking Skills in Engineering Students Through AI

Amol Harsh, Brainerd Prince, Siddharth Siddharth, Deepan Raj Prabakar Muthirayan, Kabir S Bhalla, Esraaj Sarkar Gupta, Siddharth Sahu

机构 * Center for Thinking, Language and Communication(思考、语言与交流中心) Plaksha University(普拉克斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22780 2025-11-10 cs.AI cs.CL cs.HC 62%

How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations

Zora Zhiruo Wang, Yijia Shao, Omar Shaikh, Daniel Fried, Graham Neubig, Diyi Yang

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05325 2025-11-10 cs.LG 57%

Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval

Janet Jenq, Hongda Shen

机构 * PitchBook USA(PitchBook美国公司)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06980 2025-11-10 cs.CL 57%

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi, Sung-Eun Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Published in IEEE Access

Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05304 2025-11-10 cs.HC 50%

psiUnity: A Platform for Multimodal Data-Driven XR

Akhil Ajikumar, Sahil Mayenkar, Steven Yoo, Sakib Reza, Mohsen Moghaddam

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05057 2025-11-10 cs.CV 50%

Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach

Yuanxiang Huangfu, Chaochao Wang, Weilei Wang

机构 * PatSnap Co., LTD.(PatSnap公司)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏