arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-19 至 2025-08-19 共收录 14 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 14 篇

2508.12803 2025-08-19 cs.CL 79%

When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

Ahmed Elshabrawy, Hour Kaing, Haiyue Song, Alham Fikri Aji, Hideki Tanaka, Masao Utiyama, Raj Dabre

机构 * MBZUAI(马克斯·普朗克人工智能研究所) NICT, Japan(日本信息通信技术研究所) IIT Madras(印度理工学院Madras分校)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12422 2025-08-19 cs.CV 67%

Illusions in Humans and AI: How Visual Perception Aligns and Diverges

Jianyi Yang, Junyi Ye, Ankan Dash, Guiling Wang

专题命中 其他安全 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01911 2025-08-19 cs.AI cs.CL cs.HC physics.comp-ph 62%

Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning

Yinggan Xu, Hana Kimlee, Yijia Xiao, Di Luo

机构 * NSF Center for Quantum Network(NSF量子网络中心) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICML 2025 Workshop on MAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06141 2025-08-19 cs.LG cs.AI 62%

Emergent Symbol-like Number Variables in Artificial Neural Networks

Satchel Grant, Noah D. Goodman, James L. McClelland

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref Transactions on Machine Learning Research (TMLR) 2835-8856 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12872 2025-08-19 cs.DB cs.CY 57%

Evaluating the Quality of Open Building Datasets for Mapping Urban Inequality: A Comparative Analysis Across 5 Cities

Franz Okyere, Meng Lu, Ansgar Brunn

专题命中 其他安全 :alignment(abstract);分类 cs.CY

Comments 25 pages, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10935 2025-08-19 cs.CV cs.LG cs.RO 57%

HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model

Qi Liu, Yabei Li, Hongsong Wang, Lei He

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12488 2025-08-19 cs.HC cs.AI 57%

Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing Process

Mohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry, Joseph Jay Williams, Anastasia Kuzminykh

机构 * University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Journal ref PACMHCI (CSCW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11280 2025-08-19 cs.CL 57%

High-Dimensional Interlingual Representations of Large Language Models

Bryan Wilie, Samuel Cahyawijaya, Junxian He, Pascale Fung

机构 * Hong Kong University of Science and Technology(香港理工大学) Cohere

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03885 2025-08-19 cs.LG 57%

Seldonian Reinforcement Learning for Ad Hoc Teamwork

Edoardo Zorzi, Alberto Castellini, Leonidas Bakopoulos, Georgios Chalkiadakis, Alessandro Farinelli

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Presented at the 2nd Reinforcement Learning Conference (RLC2025), Edmonton, Canada. To be published in the Proceedings of the Reinforcement Learning Journal 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11889 2025-08-19 cs.CL 57%

In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning

Hui Ma, Bo Zhang, Jinpeng Hu, Zenglin Shi

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12347 2025-08-19 cs.AR 50%

An ECC-based Fault Tolerance Approach for DNNs

Mohsen Raji, Mohammad Zaree, Kimia Soroush

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12137 2025-08-19 cs.CV 50%

Infusing fine-grained visual knowledge to Vision-Language Models

Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo, Ondřej Chum

机构 * VRG, FEE, Czech Technical University in Prague(捷克布拉格技术大学)

专题命中 其他安全 :alignment(abstract)

Comments ICCVW 2025 accepted paper. Workshop name: "What is Next in Multimodal Foundation Models?"

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09693 2025-08-19 cs.CV 50%

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

Jiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang, Jianpeng Chen, Xusen Hei, Jiayuan Xie, Yi Cai, Qing Li

机构 * South China University of Technology(南方科技大学) The Hong Kong Polytechnic University(香港理工大学) Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室)

专题命中 其他安全 :safety(abstract)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04354 2025-08-19 cs.IT eess.SP math.AG math.IT 50%

A transversality theorem for semi-algebraic sets with application to signal recovery from the second moment and cryo-EM

Tamir Bendory, Nadav Dym, Dan Edidin, Arun Suresh

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏