arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-10 至 2025-12-10 共收录 36 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 6 篇

2512.08329 2025-12-10 cs.CV cs.AI cs.LG 62%

Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models

在扩散模型图像保护方法中解释结构扰动

Michael R. Martin, Garrick Chan, Kwan-Liu Ma

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过分析扩散模型图像保护机制的结构化扰动,揭示其在特征层面的变形特性,为防御和检测策略设计提供新视角。

Comments 32 pages, 17 figures, 1 table, 5 algorithms, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04622 2025-12-10 cs.LG cs.AI cs.NE 62%

Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families

测量度量:跨模型家族的表征相似性度量的判别能力

Jialin Wu, Shreya Saha, Yiqing Bo, Meenakshi Khosla

机构 * Department of Computer Science and Engineering, UC San Diego(计算机科学与工程系,UC San Diego) Department of Electrical and Computer Engineering, UC San Diego(电气与计算机工程系,UC San Diego) Department of Cognitive Science, UC San Diego(认知科学系,UC San Diego)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过系统评估不同表征相似性度量的判别能力,揭示了软匹配在跨模型家族分离中的最优表现,为大规模模型和脑比较提供了度量选择依据。

Comments update camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08524 2025-12-10 cs.CV cs.CL 57%

Beyond Real Weights: Hypercomplex Representations for Stable Quantization

超越真实权重:用于稳定量化 的超复数表示

Jawad Ibn Ahad, Maisha Rahman, Amrijit Biswas, Muhammad Rafsan Kabir, Robin Krambroeckers, Sifat Momen, Nabeel Mohammed, Shafin Rahman

机构 * Artificial Intelligence Department, RobotBulls Labs(机器人bulls实验室人工智能部门) Machine Intelligence Lab (MILab), North South University(北南大学机器智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出了一种基于超复数乘法的渐进式重新参数化策略,用于压缩多模态语言模型,实现参数和计算量的显著减少,同时保持模型性能。

Comments Accepted in Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00218 2025-12-10 cs.AI cs.CR 57%

Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?

压力下的推理:训练激励如何影响推理链的可监控性?

Matt MacDermott, Qiyao Wei, Rada Djoneva, Francis Rhys Ward

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文研究了训练激励对推理链可监控性的影响,发现对抗性优化降低监控性能,而直接优化可监控性未显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22884 2025-12-10 cs.DC cs.AI cs.ET cs.NI cs.SY eess.SY 57%

Performance Measurements in the AI-Centric Computing Continuum Systems

面向AI导向计算连续体系统的性能测量

Praveen Kumar Donta, Qiyang Zhang, Schahram Dustdar

机构 * Department of Computer Systems and Sciences(计算机系统科学系) Stockholm University(斯德哥尔摩大学) Computer Science School(计算机科学学院) Peking University(北京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文探讨了AI导向计算连续体系统中性能测量的挑战与方法,提出新的性能维度以适应不断变化的计算需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08935 2025-12-10 cs.IR 50%

Personalize Before Retrieve: LLM-based Personalized Query Expansion for User-Centric Retrieval

在检索前个性化:基于LLM的个性化查询扩展用于以用户为中心的检索

Yingyi Zhang, Pengyue Jia, Derong Xu, Yi Wen, Xianneng Li, Yichao Wang, Wenlin Zhang, Xiaopeng Li, Weinan Gan, Huifeng Guo, Yong Liu, Xiangyu Zhao

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出PBR框架,通过在检索前个性化查询扩展,提升用户个性化检索的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏