arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-03 至 2026-03-03 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 隐私与版权 6 篇

2603.00061 2026-03-03 cs.CR cs.LG 79%

The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage

领域微调的隐性成本:包含个人身份信息的数据会削弱安全性和增加泄露

Jayesh Choudhari, Piyush Kumar Singh

专题命中 隐私与版权 :safety(title,abstract);分类 cs.LG

AI总结 研究揭示领域微调导致安全性和隐私泄露风险增加,包含PII数据加剧了有害合规和信息泄露问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05608 2026-03-03 cs.CR cs.AI cs.CL cs.LG 69%

BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints

BinaryShield: 在LLM服务中利用隐私保护指纹实现跨服务威胁情报

Waris Gill, Natalie Isak, Matthew Dressman

机构 * Microsoft Redmond, USA(微软红mond分校) Microsoft New York, USA(微软纽约分校)

专题命中 隐私与版权 :prompt injection(abstract);分类 cs.CL、cs.AI、cs.LG;trustworthy(comments)

AI总结 BinaryShield通过隐私保护技术实现跨服务威胁情报共享,提升LLM安全防护能力。

Comments Accepted at the 2026 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07940 2026-03-03 cs.CV cs.AI cs.CL cs.LG cs.MM 67%

TTOM: Test-Time Optimization and Memorization for Compositional Video Generation

TTOM:测试时优化与记忆化用于组合视频生成

Leigang Qu, Ziyang Wang, Na Zheng, Wenjie Wang, Liqiang Nie, Tat-Seng Chua

机构 * National University of Singapore(国立新加坡大学) University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 TTOM通过测试时优化与记忆化机制,提升组合视频生成的跨模态对齐能力,实现高效且可扩展的实时生成。

Comments ICLR 2026 Camera-ready. Project page: https://ttom-t2v.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04403 2026-03-03 cs.CY cs.AI cs.HC 62%

Balancing Usability and Compliance in AI Smart Devices: A Privacy-by-Design Audit of Google Home, Alexa, and Siri

在AI智能设备中平衡易用性与合规性:对Google Home、Alexa和Siri的隐私设计审计

Trevor De Clark, Yulia Bobkova, Ajay Kumar Shrestha

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文通过隐私设计审计评估Google Home、Alexa和Siri的隐私与易用性,发现易用性与合规性之间存在权衡,需提升透明度和政策一致性以保护年轻用户隐私。

Comments Published in the 2026 IEEE CCWC proceedings

Journal ref 2026 IEEE 16th Annual Computing and Communication Workshop and Conference (CCWC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00194 2026-03-03 cs.CV cs.AI cs.CR 57%

SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models

SKeDA:一种面向文本到视频扩散模型的生成水印框架

Yang Yang, Xinze Zou, Zehua Ma, Han Fang, Weiming Zhang

机构 * School of Electronic and Information Engineering, Anhui University(安徽大学电子与信息工程学院) Anhui Province Key Laboratory of Digital Security and the CAS Key Laboratory of Electromagnetic Space Information, University of Science and Technology of China(安徽省数字安全重点实验室和中国科学院电磁空间信息重点实验室,中国科学技术大学) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI

AI总结 SKeDA是一种专为文本到视频扩散模型设计的生成水印框架,通过分布保持采样和差分注意机制提升水印的鲁棒性和可靠性。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00081 2026-03-03 cs.CY 57%

From Framework to Practice: Youth Negotiations of Privacy with Smart Voice Assistants Through the PEA-AI Lens

从框架到实践:通过PEA-AI视角探讨青少年与智能语音助手的隐私协商

Molly Campbell, Yulia Bobkova, Ajay Kumar Shrestha

专题命中 隐私与版权 :alignment(abstract);分类 cs.CY

AI总结 本研究通过PEA-AI框架探讨青少年与智能语音助手的隐私协商,揭示隐私悖论并提出设计原则与治理建议。

Comments submitted to the ACM journal; "in review"

详情

展开后加载摘要…

URL PDF HTML 收藏