arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-17 至 2025-10-17 共收录 49 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 9 篇

2510.13931 2025-10-17 cs.CL 79%

Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions

Siying Liu, Shisheng Zhang, Indu Bala

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL

Comments Preprint of a paper accepted as a poster at the NeurIPS 2025 Workshop on Generative AI for Health (GenAI4Health). The final camera-ready workshop version may differ. Licensed under CC BY 4.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07887 2025-10-17 cs.CL cs.AI 73%

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia

机构 * University of Calabria(卡利博利亚大学)

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Journal ref Cantini, R., Orsino, A., Ruggiero, M., Talia, D. Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judge. Mach Learn 114, 249 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14053 2025-10-17 cs.AI 70%

Position: Require Frontier AI Labs To Release Small "Analog" Models

Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer, Philip Quirke

机构 * Frontier AI Labs(前沿人工智能实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14106 2025-10-17 cs.AI cs.CL cs.GT 62%

Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Carter Blair, Kate Larson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08236 2025-10-17 cs.LG cs.AI 62%

The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models

Konrad Löhr, Shuzhou Yuan, Michael Färber

机构 * Technische Universität Dresden(德累斯顿技术大学) Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能研究中心(ScaDS.AI))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14443 2025-10-17 cs.SD cs.AI eess.AS 57%

Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare

Mayuri Kate, Suresh Neethirajan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 40 pages, 14 figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10623 2025-10-17 cs.LG cs.CV 57%

Flows and Diffusions on the Neural Manifold

Daniel Saragih, Deyu Cao, Tejas Balaji

机构 * Queen’s University and Vector Institute(女王大学和向量研究所) University of Tokyo(东京大学) University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 43 pages, 11 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13699 2025-10-17 cs.IR cs.AI 57%

A Comprehensive Review of Recommender Systems: Transitioning from Theory to Practice

Shaina Raza, Mizanur Rahman, Safiullah Kamawal, Armin Toroghi, Ananya Raval, Farshad Navah, Amirmohammad Kazemeini

机构 * Vector Institute(向量研究所) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments we quarterly update of this literature

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他安全 11 篇

2502.11401 2025-10-17 cs.CL 79%

Following the Autoregressive Nature of LLM Embeddings via Compression and Alignment

Jingcheng Deng, Zhongtao Jiang, Liang Pang, Liwei Chen, Kun Xu, Zihao Wei, Huawei Shen, Xueqi Cheng

机构 * Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室,计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) Kuaishou Technology(快手科技)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15693 2025-10-17 cs.CV cs.MM 78%

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

Cristian Sbrolli, Matteo Matteucci

机构 * Department of Electronics, Information and Bioengineering(电子、信息与生物工程系)

专题命中 其他安全 :alignment(title,abstract)

Comments to appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17692 2025-10-17 cs.CL cs.AI cs.LG 67%

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang

机构 * Beihang University(北航) M-A-P The Hong Kong Polytechnic University(香港理工大学) AIWaves University of Alberta(阿尔伯塔大学) University of Waterloo(滑铁卢大学) University of Manchester(曼彻斯特大学) Chinese Academy of Sciences(中国科学院) Peking University(北京大学) Shanghai AI Lab(上海AI实验室) Nanjing University(南京大学) Kuaishou Technology(快手科技)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025 (Oral). Codes and models are available in https://github.com/MIO-Team/MIO

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14620 2025-10-17 cs.CL cs.AI 62%

Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models

Kedi Chen, Zhikai Lei, Xu Guo, Xuecheng Wu, Siyuan Zeng, Jianghao Yin, Yinqi Zhang, Qin Chen, Jie Zhou, Liang He, Qipeng Guo, Kai Chen, Wei Zhang

机构 * East China Normal University(东华大学) Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Xi’an Jiaotong University(西安交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14727 2025-10-17 cs.LG 57%

The Pursuit of Diversity: Multi-Objective Testing of Deep Reinforcement Learning Agents

Antony Bartlett, Cynthia Liem, Annibale Panichella

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Pre-print - Accepted at Symposium on Search Based Software Engineering (SSBSE) 2025 co-located with ASE'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14702 2025-10-17 cs.AI 57%

Cognitive-Aligned Spatio-Temporal Large Language Models For Next Point-of-Interest Prediction

Penglong Zhai, Jie Li, Fanyi Di, Yue Liu, Yifang Yuan, Jie Huang, Peng Wu, Sicong Wang, Mingyang Yin, Tingting Hu, Yao Xu, Xin Li

机构 * AMAP, Alibaba Group(阿里集团AMAP)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14387 2025-10-17 cs.AI 57%

Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?

Yijie Hu, Zihao Zhou, Kaizhu Huang, Xiaowei Huang, Qiufeng Wang

机构 * Xi’an-Jiaotong Liverpool University(西交利物浦大学) University of Liverpool(利物浦大学) Duke Kunshan University(杜克昆山大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14223 2025-10-17 cs.IR cs.AI 57%

Large Scale Retrieval for the LinkedIn Feed using Causal Language Models

Sudarshan Srinivasa Ramanujam, Antonio Alonso, Saurabh Kataria, Siddharth Dangi, Akhilesh Gupta, Birjodh Singh Tiwana, Manas Somaiya, Luke Simon, David Byrne, Sojeong Ha, Sen Zhou, Andrei Akterskii, Zhanglong Liu, Samira Sriram, Crescent Xiong, Zhoutao Pei, Angela Shao, Alex Li, Annie Xiao, Caitlin Kolb, Thomas Kistler, Zach Moore, Hamed Firooz

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12498 2025-10-17 cs.CY cs.HC 57%

The Start Button Problem: a basis for human responsibility in artificial intelligence computation

Vincenzo Calderonio

专题命中 其他安全 :alignment(abstract);分类 cs.CY

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03479 2025-10-17 cs.CL 57%

Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi

Sourabrata Mukherjee, Atharva Mehta, Sougata Saha, Akhil Arora, Monojit Choudhury

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted and published at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13821 2025-10-17 cs.NI 50%

LLM Agent Communication Protocol (LACP) Requires Urgent Standardization: A Telecom-Inspired Protocol is Necessary

Xin Li, Mengbing Liu, Chau Yuen

专题命中 其他安全 :safety(abstract)

Comments Accepted at NeurIPS 2025 AI4NextG Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏