arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-21 至 2025-11-21 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2511.16402 2025-11-21 cs.AI cs.DB 80%

Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance

可信AI在代理湖仓中的实现:从并发到治理

Jacopo Tagliabue, Federico Bianchi, Ciro Greco

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文提出Bauplan设计,通过事务机制实现湖仓中的数据和计算隔离,解决代理工作流的可信性问题,并提供自修复管道的实现。

Comments AAAI26, pre-print of paper accepted at the Trustworthy Agentic AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24143 2025-11-21 cs.DC 78%

Enhancing Traffic Safety with AI and 6G: Latency Requirements and Real-Time Threat Detection

用AI和6G提升交通安全性:延迟需求与实时威胁检测

Kurt Horvath, Dragi Kimovski, Stojan Kitanov, Radu Prodan

专题命中 安全评测 :safety(title,abstract)

AI总结 本文提出基于6G和AI的交通安全框架,通过实时威胁检测和低延迟通信提升交通安全性。

Comments Sumbitted/Accepted ICINT 2025 (PrePrint)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15716 2025-11-21 cs.AI 70%

MACIE: Multi-Agent Causal Intelligence Explainer for Collective Behavior Understanding

MACIE:多智能体因果智能解释器用于集体行为理解

Abraham Itzhak Weinberg

机构 * AI-WEINBERG, AI Experts, Tel Aviv, Israel(AI-WEINBERG,AI专家,特拉维夫,以色列)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 MACIE通过结合因果模型和反事实分析,为多智能体强化学习提供全面的因果解释,解决集体行为归因、涌现智能量化和可操作解释问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16410 2025-11-21 cs.SE 67%

Data Annotation Quality Problems in AI-Enabled Perception System Development

人工智能感知系统开发中的数据标注质量问题

Hina Saeeda, Tommy Johansson, Mazen Mohamad, Eric Knauss

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 本研究通过多组织案例研究,提出了一种涵盖完整性、准确性和一致性的18种标注错误类型分类法,为人工智能感知系统开发中的数据标注质量提供了共享词汇、诊断工具和可操作指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16482 2025-11-21 cs.LG cs.AI stat.ML 62%

Correlation-Aware Feature Attribution Based Explainable AI

基于相关性的特征归因可解释人工智能

Poushali Sengupta, Yan Zhang, Frank Eliassen, Sabita Maharjan

机构 * Department of Informatics, University of Oslo, Norway(奥斯陆大学信息学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 ExCIR通过相关性感知的归因方法,提供高效、一致且可扩展的可解释人工智能解决方案。

Comments Accepted, 2026 International Conference on Advances in Artificial Intelligence and Machine Learning (AAIML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00091 2025-11-21 cs.CY cs.AI 62%

A First-Principles Based Risk Assessment Framework and the IEEE P3396 Standard

基于第一性原理的风险评估框架及IEEE P3396标准

Richard J. Tong, Marina Cortês, Jeanine A. DeFalco, Mark Underwood, Janusz Zalewski

机构 * Chair, IEEE Artificial Intelligence Standards Committee (AISC) Institute of Astrophysics Space Sciences, University of Lisbon, Portugal Vice Chair, IEEE Artificial Intelligence Standards Committee University of New Haven Florida Gulf Coast University, United States State Academy of Applied Sciences, Ciechanow, Poland

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文提出基于第一性原理的生成式AI风险评估框架,通过信息分类系统识别不同风险类型并归责于相关方,旨在提升AI治理的严谨性与责任性。

Comments 8 pages with 3 tables. This manuscript is prepared for publication by the Institute of Electrical and Electronics Engineers, Standards Association (IEEE-SA), Sponsor Committee - Artificial Intelligence Standards Committee (C/AISC) as a White Paper of Working Group p3396 at https://standards.ieee.org/ieee/3396/11379/

Journal ref 2025 IEEE Conference on Artificial Intelligence (CAI), pp. 1588-1595, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16438 2025-11-21 cs.CL cs.IR 57%

ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports

ESGBench:用于企业可持续发展报告中可解释ESG问答的基准

Sherine George, Nithish Saji

机构 * BNY FedEx

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 ESGBench是一个用于评估企业可持续发展报告中可解释ESG问答系统性能的基准,通过领域相关问题和人工整理答案来评估模型推理能力,揭示了事实一致性、可追溯性和领域对齐等关键挑战。

Comments Workshop paper accepted at AI4DF 2025 (part of ACM ICAIF 2025). 3 pages including tables and figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17936 2025-11-21 cs.HC cs.AI 57%

When concept-based XAI is imprecise: Do people distinguish between generalisations and misrepresentations?

基于概念的XAI有时不够精确:人们能否区分泛化与误表?

Romy Müller

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 研究探讨了人们在基于概念的XAI中区分泛化与误表的能力,发现人们更敏感于相关特征的不精确性,而非泛化不相关特征的不精确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16221 2025-11-21 cs.CV cs.CL 57%

Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions

大语言模型能读取环境吗?一个多模态基准用于评估多当事人社交互动中的欺骗

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Ruicong Liu, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 本文提出多模态互动欺骗评估任务,通过新颖数据集评估多种MLLMs的欺骗检测能力,揭示其在多模态社交线索处理上的不足,并提出SoCoT和DSEM模块提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16132 2025-11-21 cs.LG 57%

An Interpretability-Guided Framework for Responsible Synthetic Data Generation in Emotional Text

可解释性引导的负责任情感文本合成数据生成框架

Paula Joy B. Martinez, Jose Marie Antonio Miñoza, Sebastian C. Ibañez

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出一种基于SHAP的可解释性引导框架,用于生成情感文本的合成数据,提升少数情绪类别的分类性能,同时揭示合成文本在词汇丰富性和表达复杂性上的局限性。

Journal ref Association for the Advancement of Artificial Intelligence (2026). Shaping Responsible Synthetic Data in the Era of Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15870 2025-11-21 cs.CE cs.AI 57%

AquaSentinel: Next-Generation AI System Integrating Sensor Networks for Urban Underground Water Pipeline Anomaly Detection via Collaborative MoE-LLM Agent Architecture

AquaSentinel:下一代集成传感器网络的AI系统,通过协作MoE-LLM代理架构实现城市地下供水管道异常检测

Qiming Guo, Bishal Khatri, Wenbo Sun, Jinwen Tang, Hua Zhang, Wenlu Wang

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 AquaSentinel通过协作MoE-LLM代理架构实现城市地下供水管道异常检测,结合物理信息和稀疏传感技术,以低成本实现高精度的实时监测。

Comments 7 pages, 1 figure, 2 tables, Accepted to the 40th AAAI Conference on Artificial Intelligence (AAAI 2026), IAAI Deployed Applications Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15720 2025-11-21 cs.AI 57%

Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models

利用大语言和视觉-语言模型进行施工工地自动危险检测

Islem Sahraoui

机构 * University of Houston Cullen College of Engineering Department of Civil and Environmental Engineering(德克萨斯大学休斯顿分校库伦工程学院土木与环境工程系)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本研究利用大语言和视觉-语言模型,通过分析文本和图像数据,提高施工工地的安全隐患检测效率。

Comments Master thesis, University of Houton

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09109 2025-11-21 cs.CV cs.CL 57%

CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation

CAIRe:通过检索增强评估进行图像文化归因

Arnav Yayavaram, Siddharth Yayavaram, Simran Khanuja, Michael Saxon, Graham Neubig

机构 * BITS Pilani(比斯·皮兰大学) Carnegie Mellon University(卡内基梅隆大学) University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 CAIRe通过检索增强评估方法,评估图像在不同文化标签下的相关性,有效衡量文化偏见,提升跨文化公平性。

Comments Preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16369 2025-11-21 eess.SP cs.NI 50%

Reasoning Meets Representation: Envisioning Neuro-Symbolic Wireless Foundation Models

推理与表征的结合:展望神经符号无线基础模型

Jaron Fontaine, Mohammad Cheraghinia, John Strassner, Adnan Shahid, Eli De Poorter

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出神经符号无线基础模型,结合神经网络与符号推理,以解决无线通信中的可解释性、鲁棒性和合规性问题,推动6G网络的智能化发展。

Comments Accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: AI and ML for Next-Generation Wireless Communications and Networking (AI4NextG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12197 2025-11-21 cs.CV 50%

Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI

超越补丁:挖掘可解释的部分原型用于可解释AI

Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz, Maguelonne Heritier, Eric Granger

专题命中 安全评测 :alignment(abstract)

AI总结 PCMNet通过挖掘可解释的部分原型提升AI系统的可解释性、稳定性和鲁棒性,为构建可靠且一致的AI系统提供新方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16150 2025-11-21 cs.CV 50%

Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval

基于推理的嵌入:利用多模态大语言模型推理提升多模态检索

Chunxu Liu, Jiyuan Yang, Ruopeng Gao, Yuhan Zhu, Feng Zhu, Rui Zhao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Sensetime Research(商汤科技研究院) Beijing Institute of Technology(北京理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出基于推理的嵌入方法,利用多模态大语言模型的推理能力提升多模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏