arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-06 至 2026-01-06 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2601.00927 2026-01-06 cs.SI cs.AI cs.CL 62%

Measuring Social Media Polarization Using Large Language Models and Heuristic Rules

利用大语言模型和启发规则测量社交媒体极化

Jawad Chowdhury, Rezaur Rashid, Gabriel Terejanu

机构 * Department of Computer Science, University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出利用大语言模型和启发规则分析社交媒体讨论中的情感极化现象,揭示事件驱动的极化模式并提供可扩展的量化方法。

Comments Foundations and Applications of Big Data Analytics (FAB), Niagara Falls, Canada, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00816 2026-01-06 cs.AI cs.CR cs.LG 62%

MathLedger: A Verifiable Learning Substrate with Ledger-Attested Feedback

MathLedger: 一种具有账本证明反馈的可验证学习基础

Ismail Ahmad Abdullah

机构 * CNU(中国矿业大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 MathLedger通过整合形式验证、密码学证明和学习动态,提供一种可验证学习的基础,实现可审计的机器认知系统。

Comments 14 pages, 1 figure, 2 tables, 2 appendices with full proofs. Documents v0.9.4-pilot-audit-hardened audit surface with fail-closed governance, canonical JSON hashing, and artifact classification. Phase I infrastructure validation; no capability claims

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05887 2026-01-06 cs.AI 57%

A three-Level Framework for LLM-Enhanced eXplainable AI: From technical explanations to natural language

一个三层框架用于LLM增强的可解释AI:从技术解释到自然语言

Marilyn Bello, Rafael Bello, Maria-Matilde García, Ann Nowé, Iván Sevillano-García, Francisco Herrera

机构 * Andalusian Research Institute in Data Science and Computational Intelligence(数据科学与计算智能安达卢西亚研究机构) Universidad de Granada(格拉纳达大学) Department of Computer Science(计算机科学系) Artificial Intelligence Lab(人工智能实验室) Vrije Universiteit Brussel(布鲁塞尔自由大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出一个三层框架,利用大型语言模型提升AI解释的可解释性,通过动态对话解释增强社会透明度和用户信任。

Comments 22 pages, 5 figures

Journal ref Bello, M., Bello, R., García, M. M., Nowé, A., Sevillano-García, I., & Herrera, F. (2025). A Three-level Framework for LLM-enhanced Explainable AI: From Technical Explanations to Natural Language. Information Systems Frontiers, 1-22

详情

展开后加载摘要…

URL PDF HTML 收藏