arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-10 至 2025-11-10 共收录 36 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8 篇

2511.04995 2025-11-10 cs.HC cs.AI cs.CL 62%

Enhancing Public Speaking Skills in Engineering Students Through AI

Amol Harsh, Brainerd Prince, Siddharth Siddharth, Deepan Raj Prabakar Muthirayan, Kabir S Bhalla, Esraaj Sarkar Gupta, Siddharth Sahu

机构 * Center for Thinking, Language and Communication(思考、语言与交流中心) Plaksha University(普拉克斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22780 2025-11-10 cs.AI cs.CL cs.HC 62%

How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations

Zora Zhiruo Wang, Yijia Shao, Omar Shaikh, Daniel Fried, Graham Neubig, Diyi Yang

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05325 2025-11-10 cs.LG 57%

Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval

Janet Jenq, Hongda Shen

机构 * PitchBook USA(PitchBook美国公司)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06980 2025-11-10 cs.CL 57%

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi, Sung-Eun Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Published in IEEE Access

Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05304 2025-11-10 cs.HC 50%

psiUnity: A Platform for Multimodal Data-Driven XR

Akhil Ajikumar, Sahil Mayenkar, Steven Yoo, Sakib Reza, Mohsen Moghaddam

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05057 2025-11-10 cs.CV 50%

Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach

Yuanxiang Huangfu, Chaochao Wang, Weilei Wang

机构 * PatSnap Co., LTD.(PatSnap公司)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏