arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-30 至 2025-10-30 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 9 篇

2510.14205 2025-10-30 cs.CL cs.AI 81%

DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans

Bingsheng Yao, Bo Sun, Yuanzhe Dong, Yuxuan Lu, Dakuo Wang

机构 * Northeastern University(东北大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments In Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19311 2025-10-30 cs.CV cs.AI 79%

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen, Jiansheng Chen, Yuman Feng, Zhihao Xi, Diyou Liu, Kai Li, Yu Meng

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Information Network Security, People’s Public Security University of China(中国人民公安大学信息网络安全学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00814 2025-10-30 cs.CL cs.AI cs.CY 67%

Many LLMs Are More Utilitarian Than One

Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, Lav R. Varshney

机构 * Forward College(前进学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Nebraska, Lincoln(内布拉斯加大学林肯分校) Technische Universität Berlin(柏林技术大学) Humboldt Institute for Internet and Society(洪堡互联网与社会研究所) Stony Brook University(石溪大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to the Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22149 2025-10-30 cs.AI cs.LG 62%

When Truthful Representations Flip Under Deceptive Instructions?

Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng, Haotian Yu, Xiaotian Han, Pan Li

机构 * Case Western Reserve University(凯斯西储大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25150 2025-10-30 cs.CL 57%

Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR

Shreyas Gopal, Ashutosh Anshul, Haoyang Li, Yue Heng Yeo, Hexin Liu, Eng Siong Chng

机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Awarded Best Student Paper at APSIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23822 2025-10-30 cs.AI 57%

ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents

Zhenyu Zhang, Tianyi Chen, Weiran Xu, Alex Pentland, Jiaxin Pei

机构 * Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Stanford Institute for Human-Centered AI(斯坦福大学人本人工智能研究所) MIT Media Lab(麻省理工学院媒体实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25332 2025-10-30 cs.CV 50%

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu

机构 * Henan Institute of Advanced Technology, Zhengzhou University(河南高级技术研究所,郑州大学) Institute of Automation, CAS(自动化研究所,中国科学院) UCAS(中国科学院大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24937 2025-10-30 cs.HC 50%

OrchVis: Hierarchical Multi-Agent Orchestration for Human Oversight

Jieyu Zhou

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24385 2025-10-30 cs.CV 50%

When are radiology reports useful for training medical image classifiers?

Herman Bergström, Zhongqi Yue, Fredrik D. Johansson

机构 * Department of Computer Science & Engineering, Chalmers University of Technology and University of Gothenburg(计算机科学与工程系,楚姆勒斯技术大学和哥德堡大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏