arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-17 至 2026-02-17 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 5 篇

2508.08500 2026-02-17 cs.AI 79%

Large Language Models as Oracles for Ontology Alignment

大型语言模型作为本体对齐的预言机

Sviatoslav Lushnei, Dmytro Shumskyi, Severyn Shykula, Ernesto Jimenez-Ruiz, Artur d'Avila Garcez

机构 * City St George’s, University of London, UK(伦敦大学城市圣乔治学院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出利用大型语言模型作为预言机,通过验证高不确定性的对应关系子集,提升本体对齐任务的性能,在OAEI 2025中取得优异成绩。

Comments Paper accepted at the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), main conference. 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13264 2026-02-17 cs.LG cs.AI cs.CL 67%

Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models

方向性集中不确定性:一种代表方法用于生成模型的不确定性量化

Souradeep Chattopadhyay, Brendan Kennedy, Sai Munikoti, Soumik Sarkar, Karl Pazdernik

机构 * Department of Mechanical Engineering, Iowa State University, Ames, IA, USA(机械工程系,爱荷华州立大学) Pacific Northwest National Laboratory, Richland, WA, USA(太平洋西北国家实验室) Department of Statistics, North Carolina State University, Raleigh, NC, USA(统计系,北卡罗来纳州立大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出方向性集中不确定性(DCU)方法,通过基于vMF分布的嵌入集中度量化,提升生成模型的不确定性量化性能,并在多模态任务中展现良好泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14581 2026-02-17 cs.LG cs.AI 62%

Model-agnostic Selective Labeling with Provable Statistical Guarantees

模型无关的可证明统计保证选择性标注

Huipeng Huang, Wenbo Liao, Huajun Xi, Hao Zeng, Mengchen Zhao, Hongxin Wei

机构 * Department of Statistics and Data Science, Southern University of Science and Technology(统计与数据科学系,南方科技大学) Department of Mathematics, The Chinese University of HongKong(数学系,香港中文大学) School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出符合性标注方法,通过控制假发现率保证人工智能预测的可信度,从而提升大规模数据标注的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14853 2026-02-17 cs.AI 57%

Uncertainty-Aware Measurement of Scenario Suite Representativeness for Autonomous Systems

面向自主系统的场景套件代表性不确定性测量

Robab Aghazadeh Chakherlou, Siddartha Khastgir, Xingyu Zhao, Jerein Jeyachandran, Shufeng Chen

机构 * WMG, University of Warwick, Coventry, UK(沃森学院,沃里克大学,科文特里,英国)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

AI总结 本文提出了一种概率方法,通过比较场景套件与预期运营领域特征分布,量化自主系统训练数据的代表性,并利用不精确贝叶斯方法处理有限数据和不确定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13567 2026-02-17 cs.CL 57%

DistillLens: Symmetric Knowledge Distillation Through Logit Lens

DistillLens: 通过Logit Lens实现对称的知识蒸馏

Manish Dhakal, Uthman Jinadu, Anjila Budathoki, Rajshekhar Sunderraman, Yi Ding

机构 * Georgia State University(佐治亚州立大学) Auburn University(阿肯色大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

AI总结 DistillLens通过Logit Lens对称对齐学生和教师模型的思维过程,提升知识蒸馏性能。

Comments Knowledge Distillation in LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏