arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-10 至 2026-02-10 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 42 篇

2602.07259 2026-02-10 cs.AI 89%

Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective

通过战略资源分配实现激励感知的AI安全:从Stackelberg安全游戏视角出发

Cheol Woo Kim, Davin Choo, Tzeh Yuan Neoh, Milind Tambe

机构 * Harvard University(哈佛大学)

专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);alignment(abstract);分类 cs.AI

AI总结 本文提出基于Stackelberg安全游戏的AI安全框架,通过战略资源分配设计激励机制,提升AI监管的主动性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11090 2026-02-10 cs.CL cs.AI 88%

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

SafeDialBench: 一种用于多轮对话中大型语言模型安全评估的细粒度基准,针对多样化对抗攻击

Hongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng, Zhixin Bai, Zhe Cao, Meng Fang, Fan Feng, Boyan Wang, Jiaheng Liu, Tianpei Yang, Jing Huo, Yang Gao, Fanyu Meng, Xi Yang, Chao Deng, Junlan Feng

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) University of Liverpool(利物浦大学) City University of Hong Kong(香港城市大学) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) China Mobile Research Institute(中国移动研究院) China Mobile (Suzhou) Software Technology Co., Ltd(中国移动苏州软件技术有限公司)

专题命中 安全评测 :safety(title,abstract);jailbreak(title,abstract);分类 cs.CL、cs.AI

AI总结 SafeDialBench提出了一种细粒度基准,用于评估大型语言模型在多轮对话中面对多样化对抗攻击时的安全性,通过多维度分类和对抗攻击策略提升评估精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07376 2026-02-10 cs.CL 83%

Do Large Language Models Reflect Demographic Pluralism in Safety?

大语言模型在安全领域是否反映了人口多样性?

Usman Naseem, Gautam Siddharth Kashyap, Sushant Kumar Ray, Rafiq Ali, Ebad Shabbir, Abdullah Mohammad

机构 * Macquarie University(麦考瑞大学) University of Delhi(德里大学) DSEU-Okhla

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 Demo-SafetyBench通过在提示级别建模人口多样性,实现了在安全评估中兼顾可扩展性和人口鲁棒性。

Comments Accepted at EACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08369 2026-02-10 cs.AI cs.CL cs.LG 82%

MemAdapter: Fast Alignment across Agent Memory Paradigms via Generative Subgraph Retrieval

MemAdapter:通过生成子图检索实现跨代理记忆范式的快速对齐

Xin Zhang, Kailai Yang, Chenyue Li, Hao Li, Qiyu Wei, Jun'ichi Tsujii, Sophia Ananiadou

机构 * The University of Manchester(曼彻斯特大学) Stanford University(斯坦福大学) Imperial College London(伦敦帝国理工学院) National Institute of Advanced Industrial Science(国家先进工业科学与技术研究院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MemAdapter通过生成子图检索实现跨代理记忆范式的快速对齐,提升记忆检索灵活性并降低对齐成本,实验显示其性能优于现有系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05656 2026-02-10 cs.LG cs.AI 82%

Alignment Verifiability in Large Language Models: Normative Indistinguishability under Behavioral Evaluation

大语言模型对齐的可验证性:行为评估下的规范不可区分性

Igor Santos-Grueiro

机构 * Igor Santos-Grueiro(独立研究者)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了大语言模型对齐的可验证性,指出行为评估无法唯一确定潜在对齐,提出了规范不可区分性概念,并通过实验验证了在评估意识下行为基准的局限性。

Comments 10 pages. Theoretical analysis of behavioral alignment evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08136 2026-02-10 cs.CV cs.AI 81%

Robustness of Vision Language Models Against Split-Image Harmful Input Attacks

视觉语言模型对分裂图像有害输入攻击的鲁棒性

Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);safety(abstract);jailbreak(abstract)

AI总结 本研究提出分裂图像视觉陷阱攻击(SIVA),揭示视觉语言模型在面对分裂图像攻击时的安全漏洞,并通过对抗性知识蒸馏算法提升跨模型攻击效果。

Comments 22 Pages, long conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11699 2026-02-10 cs.CY 79%

Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies

前沿AI审计:迈向对领先AI公司安全与安全实践的严格第三方评估

Miles Brundage, Noemi Dreksler, Aidan Homewood, Sean McGregor, Patricia Paskov, Conrad Stosz, Girish Sastry, A. Feder Cooper, George Balston, Steven Adler, Stephen Casper, Markus Anderljung, Grace Werner, Soren Mindermann, Vasilios Mavroudis, Ben Bucknall, Charlotte Stix, Jonas Freund, Lorenzo Pacchiardi, Jose Hernandez-Orallo, Matteo Pistillo, Michael Chen, Chris Painter, Dean W. Ball, Cullen O'Keefe, Gabriel Weil, Ben Harack, Graeme Finley, Ryan Hassan, Scott Emmons, Charles Foster, Anka Reuel, Bri Treece, Yoshua Bengio, Daniel Reti, Rishi Bommasani, Cristian Trout, Ali Shahin Shamsabadi, Rajiv Dattani, Adrian Weller, Robert Trager, Jaime Sevilla, Lauren Wagner, Lisa Soder, Ketan Ramakrishnan, Henry Papadatos, Malcolm Murray, Ryan Tovcimak

专题命中 安全评测 :safety(title,abstract);分类 cs.CY

AI总结 本文提出前沿AI审计的愿景,通过四个AI保证级别提升安全与安全实践的评估标准,推动行业高质量审计生态建设。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07414 2026-02-10 cs.AI cs.CL 76%

Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution

LLMs真的能体现人类个性吗?分析AI与人类行为在纠纷解决中的对齐

Deuksin Kwon, Kaleen Shrestha, Bin Han, Spencer Lin, James Hale, Jonathan Gratch, Maja Matarić, Gale M. Lucas

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文研究LLMs在模拟人类人格驱动的冲突行为方面的有效性,通过评估框架和数据集方法,揭示LLMs在纠纷解决中与人类行为的显著差异。

Comments AAAI 2026 (Special Track: AISI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03152 2026-02-10 cs.CL cs.AI cs.LG 75%

MedVAL: Toward Expert-Level Medical Text Validation with Language Models

MedVAL: 向语言模型的专家级医学文本验证迈进

Asad Aali, Vasiliki Bikia, Maya Varma, Nicole Chiou, Sophie Ostmeier, Arnav Singhvi, Magdalini Paschali, Ashwin Kumar, Andrew Johnston, Karimar Amador-Martinez, Eduardo Juan Perez Guerrero, Paola Naovi Cruz Rivera, Sergios Gatidis, Christian Bluethgen, Eduardo Pontes Reis, Eddy D. Zandee van Rilland, Poonam Laxmappa Hosamani, Kevin R Keet, Minjoung Go, Evelyn Ling, David B. Larson, Curtis Langlotz, Roxana Daneshjou, Jason Hom, Sanmi Koyejo, Emily Alsentzer, Akshay S. Chaudhari

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MedVAL通过自监督蒸馏方法,利用合成数据提升语言模型在医学文本验证中的性能,达到专家级别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08086 2026-02-10 cs.LG 74%

Probability Hacking and the Design of Trustworthy ML for Signal Processing in C-UAS: A Scenario Based Method

概率黑客与用于无人机系统(C-UAS)信号处理的可信机器学习设计:基于场景的方法

Liisa Janssens, Laura Middeldorp

专题命中 安全评测 :trustworthy(title);分类 cs.LG

AI总结 本文提出了一种基于场景的方法,用于增强C-UAS的信号处理能力,通过识别法律机制中的要求来防止概率黑客,从而提升系统的可信度。

Comments 6 pages, Pre-publication. Copyright 2026 IEEE. Peer Reviewed. Accepted at ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), scheduled for 3-8 May 2026 in Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07963 2026-02-10 cs.CL cs.AI 73%

Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms

迷失在翻译中?复合危害的跨语言迁移比较研究

Vaibhav Shukla, Hardik Sharma, Adith N Reganti, Soham Wasmatkar, Bagesh Kumar, Vrijendra Singh

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过复合危害基准测试,探讨了跨语言迁移中安全对齐的稳定性,发现攻击成功率在印度语言中显著上升,但上下文性危害迁移较温和,强调翻译基准在多语言安全测试中的必要性。

Comments Accepted at the AICS Workshop, AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22636 2026-02-10 cs.AI 70%

Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling

在最佳-N采样下对大语言模型中的对抗风险进行统计估计

Mingqian Feng, Xiaodong Liu, Weiwei Yang, Chenliang Xu, Christopher White, Jianfeng Gao

机构 * University of Rochester(罗切斯特大学) Microsoft Research(微软研究院)

专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.AI

AI总结 本研究提出SABER方法,通过Beta分布建模样本成功概率,利用解析缩放定律预测大规模对抗风险,显著降低估计误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08672 2026-02-10 cs.CL cs.LG 62%

Learning to Judge: LLMs Designing and Applying Evaluation Rubrics

学习评判:LLMs设计与应用评价量规

Clemencia Siro, Pourya Aliannejadi, Mohammad Aliannejadi

机构 * Centrum Wiskunde & Informatica (CWI)(荷兰阿姆斯特丹数学与信息学中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文研究LLMs能否自主设计和应用评价量规,发现其在一致性与跨模型泛化上存在差异,需联合建模人类与LLM的评估语言以提升可靠性。

Comments Accepted at EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16773 2026-02-10 cs.CY cs.LG stat.AP stat.ML 62%

Advance Real-time Detection of Traffic Incidents in Highways using Vehicle Trajectory Data

利用车辆轨迹数据实现高速公路交通事故的先进实时检测

Sudipta Roy, Samiul Hasan

专题命中 安全评测 :safety(abstract);分类 cs.CY、cs.LG

AI总结 本研究利用车辆轨迹数据和机器学习算法,通过随机森林模型实现高速公路交通事故的先进实时检测,提升事故预警能力。

Comments 19 Pages, 4 Tables, 10 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07562 2026-02-10 cs.LG cs.AI stat.ML 62%

Gaussian Match-and-Copy: A Minimalist Benchmark for Studying Transformer Induction

高斯匹配与复制:一种研究Transformer归纳的极简基准

Antoine Gonon, Alexandre Cordonnier, Nicolas Boumal

机构 * Institute of Mathematics, EPFL, Switzerland(数学研究所,瑞士联邦理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出高斯匹配与复制基准,用于研究Transformer的匹配与复制机制,通过隔离长距离检索信号,分析模型在检索能力上的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08889 2026-02-10 cs.AI 57%

Scalable Delphi: Large Language Models for Structured Risk Estimation

可扩展的德尔菲:用于结构化风险估计的大型语言模型

Tobias Lorenz, Mario Fritz

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出可扩展的德尔菲方法,利用大型语言模型进行结构化风险评估,通过多样化的专家人设和迭代优化,实现了高效且准确的风险估计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08716 2026-02-10 cs.CL 57%

PERSPECTRA: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments

PERSPECTRA: 一种可扩展且可配置的多元视角论证基准

Shangrui Nie, Kian Omoomi, Lucie Flek, Zhixue Zhao, Charles Welch

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 PERSPECTRA提出了一种结合结构清晰与语义多样性的多元视角论证基准,用于评估模型在处理多个视角时的表示、区分与推理能力。

Comments 15 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08342 2026-02-10 cs.CV cs.AI 57%

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science

UrbanGraphEmbeddings: 学习和评估空间导向的多模态嵌入用于城市科学

Jie Zhang, Xingtong Yu, Yuan Fang, Rudi Stouffs, Zdravko Trivic

机构 * National University of Singapore Department of Architecture Singapore The Chinese University of Hong Kong Dept of Systems Eng. \& Eng. Mgmt. China Singapore Management University School of Computing \& Info. Systems Singapore National University of Singapore Department of Architecture The Chinese University of Hong Kong Singapore Management University

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出UGE框架,通过空间导向的多模态嵌入提升城市理解任务性能,实验显示在图像检索和地理位置排名上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08339 2026-02-10 cs.AI cs.CV 57%

CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT

CoTZero:通过分层合成CoT实现无标注的人类级视觉推理

Chengyi Du, Yazhe Niu, Dazhong Shen, Luxin Xu

机构 * University of Electronic Science and Technology of China(电子科技大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong MMLab(香港中文大学 MMLab) The College of Computer Science and Technology(计算机科学与技术学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 CoTZero通过双阶段数据合成和认知对齐训练,实现无标注的人类级视觉推理,提升模型的层次推理和泛化能力。

Comments 16 pages 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08121 2026-02-10 cs.AI 57%

Initial Risk Probing and Feasibility Testing of Glow: a Generative AI-Powered Dialectical Behavior Therapy Skills Coach for Substance Use Recovery and HIV Prevention

Glow的初始风险探测与可行性测试:一种生成式人工智能驱动的辩证行为疗法技能教练,用于物质使用康复和HIV预防

Liying Wang, Madison Lee, Yunzhang Jiang, Steven Chen, Kewei Sha, Yunhe Feng, Frank Wong, Lisa Hightow-Weidman, Weichao Yuwen

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 Glow是一种生成式人工智能驱动的辩证行为疗法技能教练,旨在通过风险探测和可行性测试评估其在物质使用康复和HIV预防中的安全性与有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10238 2026-02-10 cs.AI 57%

The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities

大语言模型的致命弱点:改变少量神经元即可摧毁语言能力

Zixuan Qin, Qingchen Yu, Kunlin Lyu, Zhaoxin Fan, Yifan Sun

机构 * Center for Applied Statistics, School of Statistics, Renmin University of China(中国人民大学统计学院应用统计中心) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(未来区块链与隐私计算北京先进创新中心) School of Artificial Intelligence, Beihang University(北航人工智能学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 研究发现大语言模型中存在关键神经元集,破坏这些神经元可导致模型崩溃,揭示了模型鲁棒性和可解释性的重要问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11472 2026-02-10 cs.CV cs.LG 57%

Toward Inherently Robust VLMs Against Visual Perception Attacks

朝向对抗视觉感知攻击的固有鲁棒性VLMs

Pedram MohajerAnsari, Amir Salarpour, Michael Kühr, Siyu Huang, Mohammad Hamad, Sebastian Steinhorst, Habeeb Olufowobi, Bing Li, Mert D. Pesé

机构 * Clemson University(克莱姆森大学) Technical Universität München(慕尼黑技术大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校)

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本文提出V2LMs,通过固有鲁棒性提升自动驾驶车辆感知对视觉攻击的抗性,实验显示其在对抗攻击下表现优于传统DNNs。

Comments Accepted to the 2026 IEEE Intelligent Vehicles Symposium (IV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00765 2026-02-10 cs.AI 57%

HouseTS: A Large-Scale, Multimodal Spatiotemporal U.S. Housing Dataset and Benchmark

HouseTS:一个大规模、多模态的美国住房时空数据集和基准

Shengkun Wang, Yanshen Sun, Fanglan Chen, Linhan Wang, Naren Ramakrishnan, Chang-Tien Lu, Yinlin Chen

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 HouseTS是一个大规模多模态美国住房时空数据集,用于提供长期房价预测的基准,涵盖超过6000个邮政编码的月度数据,并支持多种模型的评估与可解释性分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07978 2026-02-10 cs.CL 57%

Cross-Linguistic Persona-Driven Data Synthesis for Robust Multimodal Cognitive Decline Detection

跨语言人格驱动的数据合成用于鲁棒多模态认知衰退检测

Rui Feng, Zhiyao Luo, Liuyu Wu, Wei Wang, Yuting Song, Yong Liu, Kok Pin Ng, Jianqing Li, Xingyao Wang

机构 * Engineering Research Center of Intelligent Theranostics Technology and Instruments, Ministry of Education, School of Biomedical Engineering and Informatics, Nanjing Medical University(智能诊疗技术与仪器工程研究中心、教育部、生物医学工程与信息学院、南京医科大学) Institute of Biomedical Engineering, University of Oxford(生物医学工程研究所、牛津大学) Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR)(高性能计算研究所、科技研究局(A*STAR)) Department of Neurology, National Neuroscience Institute, Singapore, 308433, Singapore(神经病学部、新加坡国家神经科学研究所) Duke-NUS Medical School, Singapore, 169857, Singapore(新加坡国立大学医学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 SynCog通过跨语言人格驱动数据合成与推理链微调,提升多模态模型在认知衰退检测中的诊断性能和跨语言泛化能力。

Comments 18 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07814 2026-02-10 cs.CV cs.AI 57%

How well are open sourced AI-generated image detection models out-of-the-box: A comprehensive benchmark study

开源生成图像检测模型的即插即用性能如何:一项全面的基准研究

Simiao Ren, Yuchen Zhou, Xingyu Shen, Kidus Zewde, Tommy Duong, George Huang, Hatsanai, Tiangratanakul, Tsang, Ng, En Wei, Jiayu Xue

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文评估了16种最新图像检测模型的即插即用性能,发现无统一最佳模型,训练数据对齐影响显著,现代生成器性能优异,揭示了跨数据集泛化中的系统性故障模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07706 2026-02-10 cs.LG 57%

Dense Feature Learning via Linear Structure Preservation in Medical Data

通过线性结构保持实现密集特征学习:医学数据

Yuanyun Zhang, Mingxuan Zhang, Siyuan Li, Zihan Wang, Haoran Chen, Wenbo Zhou, Shi Li

机构 * Independent Researcher(独立研究者) The Chinese University of Hong Kong(香港中文大学) Columbia University(哥伦比亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 本文提出密集特征学习方法,通过线性结构保持提升医学数据表示的稳定性与可解释性,实现更优的下游性能。

Comments ICLR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07662 2026-02-10 cs.AI 57%

ONTrust: A Reference Ontology of Trust

ONTrust:信任的参考本体

Glenda Amaral, Tiago Prince Sales, Riccardo Baratella, Daniele Porello, Renata Guizzardi, Giancarlo Guizzardi

机构 * Semantics, Cybersecurity \& Services, University of Twente, The Netherlands Business Information Systems, University of Twente, The Netherlands DAFIST, University of Genoa, Italy

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 ONTrust通过构建信任参考本体,提供信任概念的本体学基础,支持信息建模、自动推理和语义互操作性,应用于信任管理、AI设计等领域。

Comments 46 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07652 2026-02-10 cs.CR cs.AI 57%

Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents

Agent-Fence: 在深度研究代理中映射安全漏洞

Sai Puppala, Ismail Hossain, Md Jahangir Alam, Yoonpyo Lee, Jay Yoo, Tanzim Ahad, Syed Bahauddin Alam, Sajedul Talukder

机构 * Computer Science Department, Southern Illinois University(索尔坦大学计算机科学系) Computer Science Department, University of Texas(德克萨斯大学计算机科学系) Hanyang University(翰阳大学) The Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格雷格尔工程学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 Agent-Fence通过定义14种信任边界攻击类别,评估深度代理的安全性,发现高风险类别如Denial-of-Wallet和Authorization Confusion,揭示代理在持久交互中的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22978 2026-02-10 cs.CL 57%

ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?

ShoppingComp: LLMs 是否真的准备好为您的购物车服务?

Huaixiao Tou, Ying Zeng, Yuemeng Li, Cong Ma, Muzhi Li, Minghao Li, Weijie Yuan, He Zhang, Kai Jia

机构 * ByteDance(字节跳动)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 ShoppingComp通过现实世界基准揭示LLM在购物代理任务中的局限性,强调其在信息定位、多约束验证和风险决策方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16813 2026-02-10 cs.CL 57%

Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text

认知语言身份融合评分(CLIFS):一种可扩展的认知导向方法,用于从文本中量化身份融合

Devin R. Wright, Jisun An, Yong-Yeol Ahn

机构 * Center for Complex Networks and Systems Research, Luddy School of Informatics, Computing, and Engineering, Indiana University Bloomington(复杂网络与系统研究中心,信息学、计算与工程学院,印第安纳大学布卢明顿分校) Cognitive Science Program, Indiana University Bloomington(认知科学项目,印第安纳大学布卢明顿分校) School of Data Science, University of Virginia(数据科学学院,弗吉尼亚大学) CulturePulse, Inc.(CulturePulse公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 CLIFS通过整合认知语言学与大语言模型,提供一种自动化方法来量化文本中的身份融合,有效提升暴力风险评估的准确性。

Comments Authors' accepted manuscript (postprint; camera-ready). To appear in the Proceedings of EMNLP 2025. Pagination/footer layout may differ from the Version of Record

详情

展开后加载摘要…

URL PDF HTML 收藏