arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-03 至 2025-11-03 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 6 篇

2510.27641 2025-11-03 cs.CL cs.LG cs.SY eess.SY 62%

SpecAttn: Speculating Sparse Attention

Harsh Shah

机构 * Machine Learning Department(机器学习系) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2025 Workshop on Structured Probabilistic Inference & Generative Modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23724 2025-11-03 cs.LG cs.AI 62%

SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA

Minrui Luo, Fuhang Kuang, Yu Wang, Zirui Liu, Tianxing He

机构 * Shanghai Qi Zhi Institute(上海启智研究院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Xiongan AI Institute(雄安人工智能研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27432 2025-11-03 cs.CV cs.AI 57%

Mitigating Semantic Collapse in Partially Relevant Video Retrieval

WonJun Moon, MinSeok Jung, Gilhan Park, Tae-Young Kim, Cheol-Ho Cho, Woojin Jun, Jae-Pil Heo

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accpeted to NeurIPS 2025. Code is available at https://github.com/admins97/MSC_PRVR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27576 2025-11-03 eess.SP 50%

Trends and Challenges in Next-Generation GNSS Interference Management

Leatile Marata, Mariona Jaramillo-Civill, Tales Imbiriba, Petri Välisuo, Heidi Kuusniemi, Elena Simona Lohan, Pau Closas

专题命中 其他安全 :safety(abstract)

Comments Submitted to AESM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15616 2025-11-03 math.AG math.AT math.NT 50%

Approximate Fiber Products of Schemes and Their Étale Homotopical Invariants

Dongfang Zhao

专题命中 其他安全 :alignment(abstract)

Comments Several experts pointed out technical flaws of this work, for example the incorrect notations being used in Section 3 and the weak connection to the claim LLM applications in Section 1. We think it is best to be withdrawn at this point so that readers will not be misled

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05079 2025-11-03 cs.SE 50%

LLM-Guided Scenario-based GUI Testing

Shengcheng Yu, Yuchen Ling, Chunrong Fang, Quan Zhou, Yi Zhao, Chunyang Chen, Shaomin Zhu, Zhenyu Chen

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏