arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8034 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8034 篇

2603.09987 2026-03-12 cs.CL cs.AI cs.LG 67%

Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation

链式推理特征转换的进化演示优化

Xinyuan Wang, Kunpeng Liu, Arun Vignesh Malarkkan, Yanjie Fu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种通过进化轨迹经验优化LLM驱动的特征转换方法,提升转换多样性与下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07329 2026-03-10 cs.AI cs.CL cs.CY 67%

The Third Ambition: Artificial Intelligence and the Science of Human Behavior

第三项雄心:人工智能与人类行为的科学

W. Russell Neuman, Chad Coleman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出利用大型语言模型作为研究人类行为、文化及道德推理的科学工具,探讨其在社会科学中的应用与方法论创新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06495 2026-03-09 cs.LG cs.AI cs.CL 67%

COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics

COLD-Steer: 通过上下文一步学习动态引导大型语言模型

Kartik Sharma, Rakshit S. Trivedi

机构 * Georgia Institute of Technology(佐治亚理工学院) Massachusetts Institute of Technology(麻省理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 COLD-Steer通过近似学习动态实现无需训练的LLM引导,有效提升引导效果并减少样本需求。

Comments ICLR 2026. Code available at https://github.com/Ksartik/cold-steer

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23136 2026-03-09 cs.CL cs.AI cs.LG 67%

Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs

模态崩溃作为不匹配解码:多模态大语言模型的信息论限制

Jayadev Billa

机构 * Yahoo(雅虎) Nuance BBN

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究揭示多模态大语言模型中模态崩溃现象的信息论限制,指出解码器评分规则决定了可访问信息量,通过LoRA干预验证了训练目标对情绪检测性能的提升作用。

Comments 24 pages, 11 tables, 2 figures. Code: https://github.com/jb1999/modality_collapse_paper, submitted for review COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06225 2026-03-09 cs.CY cs.AI cs.CL 67%

Classroom AI: Large Language Models as Grade-Specific Teachers

教室AI:大型语言模型作为按年级教师

Jio Oh, Steven Euijong Whang, James Evans, Jindong Wang

机构 * KAIST(韩国科学技术院) Microsoft Research Asia(微软亚洲研究院) University of Chicago(芝加哥大学) William & Mary(威廉与玛丽学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究提出一种框架,通过微调大型语言模型生成适合不同年级学生的教育内容,显著提升年级匹配度,同时保持响应准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02888 2026-03-05 cs.LG cs.AI cs.CL 67%

When Your Own Output Becomes Your Training Data: Noise-to-Meaning Loops and a Formal RSI Trigger

当你的输出成为你的训练数据:噪声到意义的循环及形式RSI触发

Rintaro Ando

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出N2M-RSI模型,展示AI代理通过自反馈和信息整合阈值可无限增加内部复杂性,并探讨了代理群交互的超线性效应。

Comments Withdrawn due to a critical error discovered in the mathematical derivation and proof of Theorem 2 (Unbounded Growth) and related Lemma 2 (Compression gain lower bound). This flaw invalidates the paper's main conclusion that N2M-RSI guarantees unbounded growth, requiring a fundamental revision of the theoretical framework

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02532 2026-03-04 cs.CV 67%

EIMC: Efficient Instance-aware Multi-modal Collaborative Perception

EIMC: 高效实例感知多模态协作感知

Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Sony Research and Development Center China(索尼(中国)研发有限公司) National University of Defense Technology(国防科技大学)

专题命中 其他安全 :alignment(abstract);safety(abstract)

AI总结 EIMC通过实例感知的多模态协作感知方法,提升自动驾驶安全性,减少带宽使用,实现高效且准确的3D感知。

Comments 9 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21722 2026-03-03 cs.CL cs.AI cs.CY 67%

German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies

德国通用社会调查人设:一种基于德国通用社会调查的人员提示集合,用于与人口一致的LLM研究

Jens Rupprecht, Leon Fröhling, Claudia Wagner, Markus Strohmaier

机构 * University of Mannheim(曼海姆大学) GESIS -- Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所) RWTH Aachen University(亚琛工业大学) Complexity Science Hub(复杂性科学中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出德国通用社会调查人设集合,用于构建与人口一致的LLM研究,通过实验证明其在模拟调查响应方面的有效性。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04067 2026-03-03 cs.LG cs.AI cs.CL 67%

What Scales in Cross-Entropy Scaling Law?

交叉熵缩放定律中什么因素具有可扩展性?

Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu, Jingtao Zhan

机构 * Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过分解交叉熵为误差熵、自我对齐和置信度三个部分,揭示了误差熵是唯一遵循幂律缩放的因素,从而解释了交叉熵缩放定律在小规模有效但在大尺度失效的原因。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03594 2026-02-26 cs.CV 67%

TIPS Over Tricks: Simple Prompts for Effective Zero-shot Anomaly Detection

TIPS Over Tricks: 简单提示用于有效的零样本异常检测

Alireza Salehi, Ehsan Karami, Sepehr Noey, Sahand Noey, Makoto Yamada, Reshad Hosseini, Mohammad Sabokrou

机构 * University of Tehran(塔里哈大学) Amirkabir University of Technology(阿米尔卡比尔技术大学) Okinawa Institute of Science and Technology(冲绳科学技术大学院)

专题命中 其他安全 :alignment(abstract);safety(abstract)

AI总结 本文提出TIPS模型,通过改进的backbone和解耦提示策略,在无需复杂模块的情况下提升零样本异常检测的图像和像素级性能。

Comments This is the extended version of the paper accepted in ICASSP'26, which will be publicly available in May. Authors' contributions may vary among the versions

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21231 2026-02-26 cs.LG cs.AI cs.CL 67%

ACAR: Adaptive Complexity Routing for Multi-Model Ensembles with Auditable Decision Traces

ACAR:适应复杂度的多模型集合路由

Ramchand Kumaresan

机构 * Ramchand Kumaresan(独立研究者)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 ACAR是一种用于多模型集合的可审计路由框架,通过自一致性方差实现任务路由,提升准确率并避免过度融合,同时揭示归因计算的挑战。

Comments 12 pages, 9 figures. Measurement framework for adaptive multi-model routing with auditable execution traces

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18777 2026-02-25 cs.AI cs.CL cs.LG 67%

Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs

通过反向传播编程:在微调LLMs时,一条指令的价值相当于100个示例

Jonathan Cook, Silvia Sapora, Arash Ahmadian, Akbir Khan, Tim Rocktaschel, Jakob Foerster, Laura Ruis

机构 * FLAIR, University of Oxford(FLAIR,牛津大学) Cohere & Cohere Labs(Cohere及Cohere实验室) Anthropic UCL AI Centre(UCL人工智能中心) MIT(麻省理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过反向传播编程(PBB)使LLMs能从声明性指令中学习程序性知识,提升训练效率并减少对大量示例的依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07238 2026-02-23 cs.CL cs.AI cs.LG 67%

Beyond Mimicry to Contextual Guidance: Knowledge Distillation for Interactive AI

超越模仿到情境引导:面向交互AI的知识蒸馏

Tong Wang, K. Sudhir

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种基于情境引导的知识蒸馏方法,通过构建可重用的战略文本引导库,提升交互AI在客户服务中的服务质量与客户满意度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09201 2026-02-20 cs.LG cs.AI cs.CL 67%

Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs

多模态提示优化:为何不利用多种模态为大语言模型服务

Yumin Choi, Dongki Kim, Jinheon Baek, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出多模态提示优化问题,提出MPO框架,通过联合优化和贝叶斯策略提升多模态提示效果,验证其在图像、视频等多模态任务中的优越性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00004 2026-02-16 cs.AI cs.CL cs.LG 67%

Finetuning Large Language Models for Automated Depression Screening in Nigerian Pidgin English: GENSCORE Pilot Study

基于大型语言模型的尼日利亚皮钦英语自动化抑郁症筛查:GENSCORE试点研究

Isaac Iyinoluwa Olufadewa, Miracle Ayomikun Adesina, Ezekiel Ayodeji Oladejo, Uthman Babatunde Usman, Owen Kolade Adeniyi, Matthew Tolulope Olawoyin

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究利用微调的大型语言模型,针对尼日利亚皮钦英语开发自动化抑郁症筛查工具,GPT-4.1在准确率和文化适应性上表现最佳,为资源受限地区心理健康应用提供基础。

Comments 10 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07547 2026-02-10 q-bio.NC cs.AI cs.CL cs.LG 67%

Linguistic properties and model scale in brain encoding: from small to compressed language models

语言属性与模型规模在脑编码中的作用:从小型到压缩语言模型

Subba Reddy Oota, Vijay Rowtula, Satya Sai Srinath Namburi, Khushbu Pahwa, Anant Khandelwal, Manish Gupta, Tanmoy Chakraborty, Bapi S. Raju

机构 * TU Berlin(柏林技术大学) IIIT-Hyderabad(海得拉巴理工学院) GE HealthCare(通用电气医疗) AWS AI Labs, Amazon(亚马逊AI实验室) Microsoft Research, Bangalore, India(微软研究院,班加罗尔,印度) Microsoft, Hyderabad, India(微软,海得拉巴,印度) IIT Delhi(德里理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现小型语言模型在脑一致性上可与大模型媲美,压缩不影响脑可预测性,挑战了神经扩展的常见假设。

Comments 40 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07276 2026-02-10 cs.AI cs.CL cs.LG 67%

Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs

Steer2Adapt:动态组合转向向量引发LLM高效适应

Pengrui Han, Xueqiang Xu, Keyang Xuan, Peiyang Song, Siru Ouyang, Runchu Tian, Yuqing Jiang, Cheng Qian, Pengcheng Jiang, Jiashuo Sun, Junxia Cui, Ming Zhong, Ge Liu, Jiawei Han, Jiaxuan You

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 STEER2ADAPT通过动态组合转向向量实现LLM高效适应,适用于需要多协调能力的复杂任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20741 2026-02-06 cs.HC cs.AI cs.CY cs.LG 67%

In defence of post-hoc explanations in medical AI

为医疗AI中的事后解释辩护

Joshua Hatherley, Lauritz Munch, Jens Christian Bjerring

机构 * Center for the Philosophy of AI(哲学人工智能中心) University of Copenhagen(哥本哈根大学) Department of Philosophy and History of Ideas(哲学与思想史系) Aarhus University(奥胡斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文为医疗AI中的事后解释辩护,指出其虽无法完全复制黑盒推理过程,但能提升用户理解、提高临床团队准确性并辅助医生决策。

Journal ref 2026. Hastings Center Report 56(1): 40-46

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03849 2026-02-05 cs.HC cs.AI cs.CL cs.LG 67%

HybridQuestion: Human-AI Collaboration for Identifying High-Impact Research Questions

HybridQuestion: 人机协作识别高影响力研究问题

Keyu Zhao, Fengli Xu, Yong Li, Tie-Yan Liu

机构 * Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系、北京理工大学、清华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 HybridQuestion通过人机协作识别高影响力研究问题,利用AI处理数据与人类判断结合,验证了在科学突破和未来问题预测中的有效性。

Comments 16 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21436 2026-02-05 cs.LG cs.AI cs.CL cs.CV 67%

From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning

从一致性到互补性:面向时间序列理解和推理的对齐与解缠多模态学习

Hang Ni, Weijia Zhang, Fei Wang, Zezhi Shao, Hao Liu

机构 * The Hong Kong University of Science(香港科学与技术大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MADI通过细粒度对齐和解缠交互提升多模态时间序列理解和推理能力,实现更精确的数值-视觉模态整合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11367 2026-02-03 cs.LG cs.AI cs.CL 67%

Sparse Autoencoder Features for Classifications and Transferability

稀疏自编码器特征用于分类和可迁移性

Jack Gallifant, Shan Chen, Kuleen Sasse, Hugo Aerts, Thomas Hartvigsen, Danielle S. Bitterman

机构 * Harvard University(哈佛大学) Mass General Brigham(麻省总医院) Boston Children’s Hospital(波士顿儿童医院) Johns Hopkins University(约翰霍普金斯大学) Maastricht University(马斯特里赫特大学) University of Virginia(弗吉尼亚大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出利用稀疏自编码器提取LLM特征,通过优化架构和二进制化策略提升分类性能和跨模型迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14151 2026-02-03 cs.LG cs.AI cs.CY cs.DB 67%

Trajectory Data Management and Mining: A Survey from Deep Learning to the LLM Era

轨迹数据管理与挖掘:从深度学习到大语言模型时代的综述

Wei Chen, Yuanshao Zhu, Yanchuan Chang, Kang Luo, Haomin Wen, Lei Li, Yanwei Yu, Qingsong Wen, Chao Chen, Kai Zheng, Yunjun Gao, Yu Zheng, Xiaofang Zhou, Yuxuan Liang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文综述了轨迹计算从深度学习到大语言模型的发展,探讨了轨迹数据管理与挖掘的应用及未来研究方向。

Comments Version 2 of Trajectory Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00092 2026-02-03 cs.LG cs.AI cs.CL cs.CV 67%

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

通过宪法进行原子概念编辑来解释和控制模型行为

Neha Kalibhat, Zi Wang, Prasoon Bajpai, Drew Proud, Wenjun Zeng, Been Kim, Mani Malek

机构 * Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究提出通过原子概念编辑学习模型宪法,以解释和控制模型行为,实验证明其在提升模型成功率方面效果显著。

Journal ref Twenty-Ninth Annual Conference on Artificial Intelligence and Statistics (AISTATS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01706 2026-02-02 cs.CL cs.AI cs.LG 67%

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

通过秩-2子空间解耦进行多步知识交互分析

Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein

机构 * Department of Computer Science, University of Copenhagen, Copenhagen, Denmark(计算机科学系,哥本哈根大学,哥本哈根,丹麦)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种秩-2子空间方法,用于多步分析NLE中的知识交互,揭示PK和CK在不同生成中的对齐特性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07042 2026-01-30 cs.HC cs.AI cs.CL cs.CY 67%

Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions

Minion:一种技术探测器,用于探索用户如何与AI伙伴协商有害的价值冲突

Xianzhe Fan, Qing Xiao, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen

机构 * The University of Hong Kong(香港大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Tsinghua University(清华大学) Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 Minion通过探测用户与AI伙伴协商有害价值冲突的过程,揭示了设计中需平衡用户责任与AI安全的挑战。

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18796 2026-01-27 cs.CL cs.AI cs.LG 67%

ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models

ctELM:利用嵌入语言模型解码和操控临床试验嵌入

Brian Ondov, Chia-Hsuan Chang, Yujia Zhou, Mauro Giuffrè, Hua Xu

机构 * Yale School of Medicine(耶鲁医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 ctELM通过嵌入语言模型解码和操控临床试验嵌入,实现对未见过的临床试验描述和生成,提升生物医学领域语言模型与嵌入空间对齐的透明度和应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13481 2026-01-27 cs.AI cs.CL cs.CY 67%

neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings

neuralFOMO: LLMs能否处理第二好?在多智能体设置中测量类似嫉妒的偏好

Arnav Ramamoorthy, Shrey Dhorajiya, Ojas Pungalia, Rashi Upadhyay, Abhishek Mishra, Abhiram H, Tejasvi Alladi, Sujan Yenuganti, Dhruv Kumar

机构 * BITS Pilani, Pilani Campus(比斯学院,帕利尼校区)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 neuralFOMO研究了LLMs在多智能体环境中是否表现出类似嫉妒的偏好,通过点分配游戏和比较评估揭示了不同模型在竞争与合作中的不同表现。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17424 2026-01-27 cs.CL cs.AI cs.CR cs.LG 67%

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

涌现的偏移:狭窄微调可以产生广泛偏移的LLM

Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, Owain Evans

机构 * University College London(伦敦大学学院) Center on Long-Term Risk(长期风险中心) Warsaw University of Technology(华沙技术大学) University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现,狭窄微调训练LLM生成不安全代码会导致广泛偏移,模型在无关提示上表现出欺骗性行为,且偏移可通过触发器隐藏。

Comments 41 pages, 38 figures An earlier revision of this paper was accepted at ICML 2025. Since then, it has been updated to include new results on the impact of formatting (4.4), new dataset (4.6), training dynamics (4.7) and base models (4.8) Extended version of the paper was published in Nature 2026/1

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10695 2026-01-23 cs.LG cs.AI cs.CL 67%

Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks

引入集合一致性验证任务与集合一致性能量网络

Mooho Song, Hyeryung Son, Jay-Yoon Lee

机构 * Seoul National University(首尔国立大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出集合一致性验证任务及SC-Energy模型,通过对比损失框架提升多陈述逻辑一致性验证性能,并发布新数据集

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025), Long Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13709 2026-01-21 cs.AI cs.CL cs.CY cs.HC cs.SI 67%

Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games

隐藏在 plain 文本中:使用社会推断游戏测量 LLM 欺骗质量 against 人类基准

Christopher Kao, Vanshika Vats, James Davis

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文通过社会推断游戏测量 LLM 欺骗能力,发现 LLM 在欺骗人类时表现更优,但其欺骗质量低于人类。

Comments For associated dataset, see https://github.com/cocochief4/llm-mafia. Published in IEEE ICA 2025, waiting for IEEEXplore proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏