arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-12 至 2026-01-12 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3 篇

2601.04285 2026-01-12 cs.AI cs.HC cs.LG cs.MA 62%

A Future Capabilities Agent for Tactical Air Traffic Control

战术空交通管制中的未来能力代理

Paul Kent, George De Ath, Martin Layton, Allen Hart, Richard Everson, Ben Carvell

机构 * University of Exeter(埃克塞特大学) NATS(英国航空交通服务提供商)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出Agent Mallard,一种基于规则的前瞻性规划代理,用于战术空交通管制,通过嵌入随机数字双胞胎实现安全冲突解决,并在简化场景中展示与专家推理一致的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05624 2026-01-12 cs.CL 57%

Text Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African Languages

西语和约鲁巴语中的文本净化:一种跨语言机器学习方法用于资源有限的非洲语言

Abayomi O. Agbeyangi

机构 * Walter Sisulu University, Buffalo City Campus, East London, South Africa(沃尔特·西苏卢大学,比夫福城校区,东伦敦,南非)

专题命中 安全训练 :safety(abstract);分类 cs.CL

AI总结 本研究提出了一种跨语言机器学习方法,用于西语和约鲁巴语的文本净化,结合可解释模型和规则编辑,实现了高准确率的有毒文本检测与重写。

Comments 26 pages, 9 figures and 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15906 2026-01-12 cs.AI 57%

Darth Vecdor: An Open-Source System for Generating Knowledge Graphs Through Large Language Model Queries

Darth Vecdor:通过大语言模型查询生成知识图谱的开源系统

Jonathan A. Handler

专题命中 安全训练 :safety(abstract);分类 cs.AI

AI总结 Darth Vecdor是一个开源系统,通过大语言模型查询生成结构化的知识图谱,旨在解决LLM响应的准确性和一致性问题,以提高医疗领域应用的安全性和效率。

Comments 17 pages, 3 figures. Changes to best of my recollection: 1) Added to Acknowledgements that Darth Vecdor software was created with the help of LLMs and many other resources, clearly noted on software GitHub site, but added here. 2) Fixed content related to the Are You Sure paper that was referenced, as the previous version incorrectly summarized the article's finding

详情

展开后加载摘要…

URL PDF HTML 收藏