arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-03 至 2025-11-03 共收录 30 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2410.05102 2025-11-03 cs.CL cs.AI cs.LG 85%

SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks

Fenia Christopoulou, Ronald Cardenas, Gerasimos Lampouras, Haitham Bou-Ammar, Jun Wang

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) University College London(伦敦大学学院)

专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 27 pages, 9 figures, 5 tables. Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23953 2025-11-03 cs.LG cs.AI cs.CL cs.CY cs.GT 84%

Representative Social Choice: From Learning Theory to AI Alignment

Tianyi Qiu

机构 * Peking University(北京大学) UC Berkeley(加州大学伯克利分校) Center for Human-Compatible AI(人类兼容人工智能中心)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Journal of Artificial Intelligence Research, in press. Best Paper at NeurIPS 2024 Pluralistic Alignment Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27254 2025-11-03 cs.CL cs.AI cs.LG 82%

Languages are Modalities: Cross-Lingual Alignment via Encoder Injection

Rajan Agarwal, Aarush Gupta

机构 * University of Waterloo(多伦多大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13487 2025-11-03 cs.CL cs.AI 62%

Detecting Prefix Bias in LLM-based Reward Models

Ashwin Kumar, Yuzi He, Aram H. Markosyan, Bobbie Chern, Imanol Arrieta-Ibarra

机构 * Washington University in St Louis(华盛顿大学圣路易斯分校) Meta Platforms, Inc.(Meta平台公司)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26913 2025-11-03 cs.DC 50%

FlowMesh: A Service Fabric for Composable LLM Workflows

Junyi Shen, Noppanat Wadlom, Lingfeng Zhou, Dequan Wang, Xu Miao, Lei Fang, Yao Lu

专题命中 偏好对齐 :RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2510.26935 2025-11-03 cs.RO cs.AI cs.CL cs.FL 81%

RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification

Yunhao Yang, Neel P. Bhatt, Pranay Samineni, Rohan Siva, Zhanyang Wang, Ufuk Topcu

专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Code and data are available at: https://repv-project.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27087 2025-11-03 cs.CL cs.CY 62%

Characterizing Selective Refusal Bias in Large Language Models

Adel Khorramrouz, Sharon Levy

机构 * Rutgers University(罗格斯大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.CY

Comments 21 pages, 12 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18779 2025-11-03 cs.CL 57%

KAT-Coder Technical Report

Zizheng Zhan, Ken Deng, Jinghui Wang, Xiaojiang Zhang, Huaixi Tang, Minglei Zhang, Zhiyi Lai, Haoyang Huang, Wen Xiang, Kun Wu, Wenhao Zhuang, Shaojie Wang, Shangpeng Yan, Kepeng Lei, Zongxian Feng, Huiming Wang, Zheng Lin, Mengtong Li, Mengfei Xie, Yinghan Cui, Xuxing Chen, Chao Wang, Weihao Li, Wenqiang Zhu, Jiarong Zhang, Jingxuan Xu, Songwei Yu, Yifan Yao, Xinping Lei, C. Zhang, Han Li, Junqi Xiong, Zuchen Gao, Dailin Li, Haimo Li, Jiaheng Liu, Yuqun Zhang, Junyi Peng, Haotian Zhang, Bin Chen

机构 * Kwaipilot Team(Kwaipilot 团队)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 5 篇

2510.27172 2025-11-03 cs.LG cs.AI 73%

Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler

Zixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) University of Central Florida(佛罗里达大学) Tsinghua University(清华大学)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27062 2025-11-03 cs.LG cs.AI 73%

Consistency Training Helps Stop Sycophancy and Jailbreaks

Alex Irpan, Alexander Matt Turner, Mark Kurzeja, David K. Elson, Rohin Shah

机构 * Google(谷歌)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26981 2025-11-03 cs.LG cs.AI 73%

Fine-Grained Iterative Adversarial Attacks with Limited Computation Budget

Zhichao Hou, Weizhi Gao, Xiaorui Liu

机构 * North Carolina State University(北卡罗来纳州立大学)

专题命中 越狱攻击 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26847 2025-11-03 cs.CR cs.AI cs.CL cs.IT math.IT 73%

Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token

Shaked Zychlinski, Yuval Kainan

机构 * JFrog

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00943 2025-11-03 cs.CR cs.AI 57%

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

Chloe Li, Mary Phuong, Noah Y. Siegel

机构 * University College London(伦敦大学学院)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI

Comments Accepted to IJCNLP-AACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2507.05538 2025-11-03 cs.AI cs.CR cs.CY 81%

Red Teaming AI Red Teaming

Subhabrata Majumdar, Brian Pendleton, Abhishek Gupta

机构 * Vijil / AI Risk and Vulnerability Alliance(Vijil / AI风险与漏洞联盟) AI Risk and Vulnerability Alliance(AI风险与漏洞联盟) Montreal AI Ethics Institute(蒙特利尔人工智能伦理研究所)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.CY

Comments Conference on Applied Machine Learning for Information Security (CAMLIS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2507.23607 2025-11-03 cs.LG cs.AI cs.CL 67%

Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates

Tien Huu Do, Antoine Masquelier, Nae Eoun Lee, Jonathan Crowther

机构 * Pfizer(辉瑞公司) Merck(默克公司)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 9 篇

2510.27077 2025-11-03 cs.CL 85%

Contrastive Knowledge Transfer and Robust Optimization for Secure Alignment of Large Language Models

Jiasen Zheng, Huajun Zhang, Xu Yan, Ran Hao, Chong Peng

专题命中 安全评测 :alignment(title,abstract);safety(abstract);trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20621 2025-11-03 cs.AI 79%

Towards the Formalization of a Trustworthy AI for Mining Interpretable Models explOiting Sophisticated Algorithms

Riccardo Guidotti, Martina Cinquini, Marta Marchiori Manerba, Mattia Setzu, Francesco Spinnato

机构 * University of Pisa(比萨大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12667 2025-11-03 cs.AI cs.LO 79%

Building Trustworthy AI by Addressing its 16+2 Desiderata with Goal-Directed Commonsense Reasoning

Alexis R. Tudor, Yankai Zeng, Huaduo Wang, Joaquin Arias, Gopal Gupta

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27207 2025-11-03 cs.LG cs.AI 62%

Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models

Hamed Najafi, Dongsheng Luo, Jason Liu

机构 * Florida International University(佛罗里达国际大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27521 2025-11-03 cs.HC cs.CY 57%

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Nick Judd, Alexandre Vaz, Kevin Paeth, Layla Inés Davis, Milena Esherick, Jason Brand, Inês Amaro, Tony Rousmaniere

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27244 2025-11-03 cs.SE cs.AI 57%

Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes

Ora Nova Fandina, Gal Amram, Eitan Farchi, Shmulik Froimovich, Raviv Gal, Wesam Ibraheem, Rami Katan, Alice Podolsky, Orna Raz

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27065 2025-11-03 cs.LG cs.PF 57%

MLPerf Automotive

Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, John Owens, Pınar Muyan-Özçelik, Tom St. John, Jinho Suh, Arjun Suresh

机构 * University of California, Davis(加州大学戴维斯分校) Arm(ARM公司) Samsung(三星) Qualcomm(高通) California State University, Sacramento(加州州立大学萨克拉门托分校) Gilmet Labs(Gilmet实验室) NVIDIA(英伟达) AMD(超微半导体)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 16 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26830 2025-11-03 cs.LG cs.CR 57%

SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation

Guangzhi Su, Shuchang Huang, Yutong Ke, Zhuohang Liu, Long Qian, Kaizhu Huang

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25761 2025-11-03 cs.CL 57%

DiagramEval: Evaluating LLM-Generated Diagrams via Graphs

Chumeng Liang, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 6 篇

2510.27641 2025-11-03 cs.CL cs.LG cs.SY eess.SY 62%

SpecAttn: Speculating Sparse Attention

Harsh Shah

机构 * Machine Learning Department(机器学习系) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2025 Workshop on Structured Probabilistic Inference & Generative Modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23724 2025-11-03 cs.LG cs.AI 62%

SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA

Minrui Luo, Fuhang Kuang, Yu Wang, Zirui Liu, Tianxing He

机构 * Shanghai Qi Zhi Institute(上海启智研究院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Xiongan AI Institute(雄安人工智能研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27432 2025-11-03 cs.CV cs.AI 57%

Mitigating Semantic Collapse in Partially Relevant Video Retrieval

WonJun Moon, MinSeok Jung, Gilhan Park, Tae-Young Kim, Cheol-Ho Cho, Woojin Jun, Jae-Pil Heo

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accpeted to NeurIPS 2025. Code is available at https://github.com/admins97/MSC_PRVR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27576 2025-11-03 eess.SP 50%

Trends and Challenges in Next-Generation GNSS Interference Management

Leatile Marata, Mariona Jaramillo-Civill, Tales Imbiriba, Petri Välisuo, Heidi Kuusniemi, Elena Simona Lohan, Pau Closas

专题命中 其他安全 :safety(abstract)

Comments Submitted to AESM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15616 2025-11-03 math.AG math.AT math.NT 50%

Approximate Fiber Products of Schemes and Their Étale Homotopical Invariants

Dongfang Zhao

专题命中 其他安全 :alignment(abstract)

Comments Several experts pointed out technical flaws of this work, for example the incorrect notations being used in Section 3 and the weak connection to the claim LLM applications in Section 1. We think it is best to be withdrawn at this point so that readers will not be misled

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05079 2025-11-03 cs.SE 50%

LLM-Guided Scenario-based GUI Testing

Shengcheng Yu, Yuchen Ling, Chunrong Fang, Quan Zhou, Yi Zhao, Chunyang Chen, Shaomin Zhu, Zhenyu Chen

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏