arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-21 至 2025-11-21 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8 篇

2511.15750 2025-11-21 cs.CY cs.AI cs.ET cs.HC 62%

Writing With Machines and Peers: Designing for Critical Engagement with Generative AI

用机器和同伴写作:设计促进生成式AI批判性参与的方法

Xinran Zhu, Cong Wang, Duane Searsmith

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出一种整合AI与同伴反馈的教学设计,通过八周的写作活动,探讨学生如何批判性地使用生成式AI并建立与AI和人类的协作关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16588 2025-11-21 cs.AI cs.LO 57%

Formal Abductive Latent Explanations for Prototype-Based Networks

基于原型网络的正式归纳性隐式解释

Jules Soria, Zakaria Chihani, Julien Girard-Satabin, Alban Grastien, Romain Xu-Darme, Daniela Cancila

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文提出基于原型网络的正式归纳性隐式解释方法,通过形式化方法提升模型可解释性,适用于图像分类任务。

Comments Accepted at AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16494 2025-11-21 cs.CV cs.AI 57%

Physics-Informed Machine Learning for Efficient Sim-to-Real Data Augmentation in Micro-Object Pose Estimation

基于物理的机器学习用于微物体位姿估计中的高效仿真到现实数据增强

Zongcai Tan, Lan Wei, Dandan Zhang

机构 * Department of Bioengineering, Imperial-X Initiative, Imperial College London(生物工程系、Imperial-X计划、帝国理工学院伦敦分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于物理的深度生成学习框架,通过整合波动光学和深度对齐,高效生成高保真显微镜图像,提升微机器人位姿估计的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04250 2025-11-21 stat.ME cs.AI cs.ET cs.IR stat.AP 57%

How many patients could we save with LLM priors?

使用 LLM 先验知识能挽救多少患者?

Shota Arai, David Selby, Andrew Vargo, Sebastian Vollmer

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 利用LLM生成的先验知识优化多中心临床试验中的不良事件建模,提高预测性能并减少患者数量需求。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17747 2025-11-21 cs.CL 57%

Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks

通过最小对ABX任务区分多语言模型中的形式与意义

Maureen de Seyssel, Jie Chi, Skyler Seto, Maartje ter Hoeve, Masha Fedzechkina, Natalie Schluter

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 通过最小对ABX任务研究多语言模型中形式与意义的区分能力,揭示了语言识别和语义识别在训练过程中的变化规律。

Comments Comments: Published in EMNLP 2025. https://aclanthology.org/2025.emnlp-main.1210.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16435 2025-11-21 cs.CV 50%

Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation

超越视觉线索:利用通用语义作为少样本分割的支持

Jin Wang, Bingfeng Zhang, Jian Pang, Mengyu Liu, Honglong Chen, Weifeng Liu

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出语言驱动属性泛化架构,通过多属性增强和多模态对齐提升少样本分割性能,实现新的最佳效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09540 2025-11-21 cs.CV 50%

vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs

vMFCoOp:在统一超球面流形上实现平衡的提示生物医学VLMs

Minye Shao, Sihan Guo, Xinrun Li, Xingyu Miao, Haoran Duan, Yang Long

专题命中 其他安全 :alignment(abstract)

AI总结 vMFCoOp通过统一超球面流形上的vMF分布估计,实现LLM与CLIP主干间的语义对齐,提升生物医学VLMs的提示效果和少样本分类性能。

Comments Accepted as an Oral Presentation at AAAI 2026 Main Technical Track (this version is not peer-reviewed; it is the extended version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15669 2025-11-21 cs.CV 50%

UINO-FSS: Unifying Representation Learning and Few-shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelation

UINO-FSS: 通过层次化蒸馏和Mamba-超相关性统一表示学习与少样本分割

Wei Zhuo, Zhiyue Tang, Wufeng Xue, Hao Ding, Junkai Ji, Linlin Shen

机构 * School of Artificial Intelligence and the National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(人工智能学院和大数据系统计算技术国家工程实验室,深圳大学) Guangdong Provincial Key Laboratory of Intelligent Information Processing, China(广东省智能信息处理重点实验室,中国) School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University(生物医学工程学院,深圳大学医学院,深圳大学) Department of Computer Science, University of Nottingham Ningbo China(计算机科学系,宁波大学中国)

专题命中 其他安全 :alignment(abstract)

AI总结 UINO-FSS通过层次化蒸馏和Mamba-超相关性整合不同基础模型知识,实现少样本分割的统一学习框架,取得新SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏