arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-17 至 2025-12-17 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2305.11581 2025-12-17 cs.AI econ.GN q-fin.EC 79%

Trustworthy, responsible, ethical AI in manufacturing and supply chains: synthesis and emerging research questions

制造业与供应链中的可信、负责任、道德AI:综合与新兴研究问题

Alexandra Brintrup, George Baryannis, Ashutosh Tiwari, Svetan Ratchev, Giovanna Martinez-Arellano, Jatinder Singh

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文探讨制造业中负责任、道德和可信AI的应用,提出研究问题以指导未来研究,确保AI应用的安全性和责任性。

Comments Pre-print under peer-review

Journal ref Data-Centric Engineering 6 (2025) e53

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13702 2025-12-17 cs.CY cs.AI 62%

Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport

提升医疗AI的透明度与可追溯性:AI产品护照

A. Anil Sinaci, Senan Postaci, Dogukan Cavdaroglu, Machteld J. Boonstra, Okan Mercan, Kerem Yilmaz, Gokce B. Laleci Erturkmen, Folkert W. Asselbergs, Karim Lekadir

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出AI产品护照,通过生命周期文档提升医疗AI的透明度和可追溯性,符合FUTURE-AI原则,确保公平性和可用性,并通过开源平台实现可访问性。

Comments A total of 33 pages: First 16 pages for the manuscript and the remaining 17 pages for the supplementary user guide of the graphical user interface

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13739 2025-12-17 cs.CV cs.AI 57%

Human-AI Collaboration Mechanism Study on AIGC Assisted Image Production for Special Coverage

面向特殊报道的AIGC辅助图像生成中的人机协作机制研究

Yajie Yang, Yuqing Zhao, Xiaochao Xi, Yinan Zhu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种人机协作机制,用于在特殊报道中通过AIGC辅助图像生成,结合高精度分割、语义对齐和风格调节技术,确保内容准确性和可验证性。

Comments AAAI-AISI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13699 2025-12-17 cs.CY 57%

Us-vs-Them bias in Large Language Models

大语言模型中的‘我们 vs 他们’偏见

Tabia Tanzin Prama, Julia Witte Zimmerman, Christopher M. Danforth, Peter Sheridan Dodds

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CY

AI总结 本研究发现大语言模型在不同人设下表现出‘我们 vs 他们’偏见,通过微调和DPO方法可有效缓解这种偏见。

详情

展开后加载摘要…

URL PDF HTML 收藏