arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

AAAI Conference on Artificial Intelligence · 会议 · Artificial Intelligence

共收录 9567
1706.07448 2026-06-04 eess.SY cs.LO cs.SY

Norm Conflict Resolution in Stochastic Domains

随机领域中的规范冲突解决

Daniel Kasenberg, Matthias Scheutz

AI总结 本文提出一种混合方法,利用线性时序逻辑在马尔可夫决策过程中管理规范冲突,同时适应领域随机性,并在模拟吸尘器领域中提供概念验证。

Comments New version of paper - new evaluations, accepted to AAAI 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08372 2026-06-04 stat.ML cs.LG cs.NA math.NA math.OC

A Unified Convex Surrogate for the Schatten-$p$ Norm

一种统一的凸替代项用于Schatten-p范数

Chen Xu, Zhouchen Lin, Hongbin Zha

AI总结 本文提出一种统一的凸替代项,用于Schatten-p范数,通过矩阵分解的等价性,使因子矩阵的范数可凸优化,提升矩阵补全任务的性能。

Comments The paper is accepted by AAAI-17. We show that multi-factor matrix factorization enjoys superiority over the traditional two-factor case

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.02314 2026-06-04 math.NA cs.NA math.OC stat.ML

A Scalable and Extensible Framework for Superposition-Structured Models

可扩展且可扩展的框架用于叠加结构模型

Shenjian Zhao, Cong Xie, Zhihua Zhang

AI总结 本文提出一种可扩展且可扩展的框架,用于解决叠加结构模型,通过近端牛顿型方法实现高效求解,并在多个数据集上展示了其强大的性能和超线性收敛速度。

Journal ref AAAI 2016: 2372-2378

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.01245 2026-06-04 math.NA cs.AI cs.NA math.OC stat.ML

Scalable Algorithms for Tractable Schatten Quasi-Norm Minimization

可扩展算法用于可计算的Schatten准范数最小化

Fanhua Shang, Yuanyuan Liu, James Cheng

AI总结 本文提出两种可计算的Schatten准范数,设计高效算法以加速大规模问题解决,并通过实验验证其精度和速度优势。

Comments 16 pages, 5 figures, Appears in Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI), Phoenix, Arizona, USA, pp. 2016--2022, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.01110 2026-06-04 math.NA cs.AI cs.LG cs.NA

Bayesian Matrix Completion via Adaptive Relaxed Spectral Regularization

基于自适应放松谱正则化的贝叶斯矩阵补全

Yang Song, Jun Zhu

AI总结 本文提出一种基于谱正则化的贝叶斯矩阵补全方法,通过放松奇异向量的正交约束,设计出适用于贝叶斯推断的自适应谱正则化方法,无需参数调优即可自动推断潜在因子数量,在稀疏矩阵上表现优异。

Comments Accepted to AAAI 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.05133 2026-06-04 math.OC cs.LG cs.NA math.NA

Fast Proximal Linearized Alternating Direction Method of Multiplier with Parallel Splitting

快速近端线性化交替方向乘子法与并行分裂

Canyi Lu, Huan Li, Zhouchen Lin, Shuicheng Yan

AI总结 本文提出快速近端增广拉格朗日法和快速近端ADMM并行分裂法,改进了收敛速度并降低了计算复杂度,实验证明其在合成和真实数据上均优于传统PALM和ADMM。

Comments AAAI 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.3536 2026-06-04 eess.SY cs.SY

A Generalized Reduced Linear Program for Markov Decision Processes

马尔可夫决策过程的通用缩减线性规划

Chandrashekar Lakshminarayanan, Shalabh Bhatnagar

AI总结 本文提出通用缩减线性规划(GRLP),通过正线性组合原始约束来解决大规模马尔可夫决策过程的计算问题,提供了新的理论框架和误差界分析。

Comments 24 pages, submitted to AAAI on November 19 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1407.1399 2026-06-04 math.NA cs.LG cs.NA

Generalized Higher-Order Tensor Decomposition via Parallel ADMM

通过并行ADMM实现广义高阶张量分解

Fanhua Shang, Yuanyuan Liu, James Cheng

AI总结 本文提出一种并行迹范数正则化的张量分解方法,通过优化方案自动确定各模式的因子数,解决传统方法在模型选择、粗腐损和计算效率上的挑战。

Comments 9 pages, 5 figures, AAAI 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1404.6871 2026-06-04 math.NA cs.CV cs.NA

Proximal Iteratively Reweighted Algorithm with Multiple Splitting for Nonconvex Sparsity Optimization

近端迭代重加权算法与多重分裂用于非凸稀疏优化

Canyi Lu, Yunchao Wei, Zhouchen Lin, Shuicheng Yan

AI总结 本文提出PIRE算法解决非凸稀疏及结构稀疏问题,相比传统方法更高效,且在每轮迭代计算成本接近凸求解器。进一步提出PIRE-PS和PIRE-AU处理多变量问题,理论证明其收敛性,实验显示性能优异。

Journal ref Twenty-Eighth AAAI Conference on Artificial Intelligence, 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07971 2026-06-03 cs.LG

Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-Tuning

低秩曲率用于大语言模型微调中的零阶优化

Hyunseok Seung, Jaewoo Lee, Hyunsuk Ko

机构 * University of Wisconsin – Madison(威斯康星大学麦迪逊分校) University of Georgia(佐治亚大学) Hanyang University(翰阳大学)

AI总结 提出LOREN方法,通过低秩块对角预条件器捕捉曲率并利用REINFORCE留一法梯度估计器降低方差,在LLM微调中实现更高精度和更快收敛,同时峰值内存使用降低27.3%。

Comments Accepted to the AAAI Conference on Artificial Intelligence (AAAI-2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00559 2026-06-02 cs.LG cs.AI

Richer Representations for Neural Algorithmic Reasoning via Auxiliary Reconstruction

通过辅助重建实现神经算法推理的更丰富表示

Jiafu Huang, Chao Peng, Chenyang Xu, Zhengfeng Yang, Kecheng Cai, Chenhao Zhang, Yi Wang, Yiwei Gong, Wanqin Zhou, Irene Zheng

机构 * sei.ecnu.edu.cn(东华大学信息科学与工程学院)

AI总结 提出辅助重建模块和自监督学习变体,增强编码器对输入状态信息的保留和特征间依赖的捕捉,从而提升神经算法推理性能。

Comments Appeared at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00003 2026-06-02 cs.CY cs.CR cs.HC

Learning from Mistakes: Can LLM Self-Recover after Misalignment?

从错误中学习:LLM 在错位后能否自我恢复?

Olga E. Sorokoletova, Francesco Giarrusso, Vincenzo Suriani, Daniele Nardi

AI总结 研究大语言模型在遭受恶意攻击后,是否具有内在的自我对齐恢复能力,并提出一种建模用户-助手交互安全轨迹并检测恢复趋势的方法。

Comments AAAI'26 Workshop (WS37), Machine Ethics: from formal methods to emergent machine ethics, January 20--27, 2026, Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22689 2026-06-02 eess.IV cs.CV

Graph-Theoretic Consistency for Robust and Topology-Aware Semi-Supervised Histopathology Segmentation

基于图论一致性的鲁棒且拓扑感知的半监督组织病理学分割

Ha-Hieu Pham, Minh Le, Han Huynh, Nguyen Quoc Khanh Le, Huy-Hieu Pham

机构 * Student(学生)

AI总结 提出拓扑图一致性(TGC)框架,通过对齐预测图与参考图的拉普拉斯谱、组件计数和邻接统计,在仅5-10%标注下实现最先进的半监督分割性能。

Comments Accepted to the AAAI 2026 Student Abstract and Poster Program

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29793 2026-05-29 cs.CV

Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

更少步骤,更优性能:基于语言的高效跨模态视频片段修剪用于视频时刻检索

Xiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou, Zichuan Xu, Wenzheng Xu, Junyang Chen, Renfu Li

机构 * Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science of Technology(湖北大数据安全工程研究中心,网络安全学院,华中科技大学) Peking University(北京大学) Henan University(河南大学) Dalian University of Technology(大连理工大学) Sichuan University(四川大学) Shenzhen University(深圳大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出SpotVMR方法,通过可学习的片段搜索模型和低成本语义索引特征,高效修剪查询相关视频片段,作为即插即用模块提升现有VMR方法的效率与性能。

Comments Published in AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12426 2026-05-29 cs.CY

Political Advertising on Facebook During the 2022 Australian Federal Election: A Social Identity Perspective

2022年澳大利亚联邦选举期间Facebook上的政治广告:社会认同视角

Stefano Civelli, Pietro Bernardelle, Frank Mols, Gianluca Demartini

AI总结 利用Meta广告库分析2022年澳大利亚联邦选举期间Facebook和Instagram上的政治广告,基于社会认同理论揭示主要政党在支出、覆盖、人口统计和地理定位上的差异,以及不同说服策略。

Journal ref Proceedings of the International AAAI Conference on Web and Social Media 20(1) (2026) 563-577

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18744 2026-05-29 cs.CL

LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex Reasoning

LogicCat:面向复杂推理的思维链文本到SQL基准测试

Tao Liu, Xutao Mao, Hongying Zan, Dixuan Zhang, Yifan Li, Haixin Liu, Lulu Kong, Jiaming Hou, Rui Li, YunLong Li, aoze zheng, Zhiqiang Zhang, Luo Zhewei, Kunli Zhang, Min Peng

机构 * Zhengzhou University(郑州大学) Vanderbilt University(范德比大学) Wuhan University(武汉大学)

AI总结 提出首个针对复杂推理和思维链解析的Text-to-SQL基准数据集LogicCat,涵盖物理、算术、常识和假设推理场景,通过4038个问题与12114条思维链步骤显著提升任务难度,现有模型执行准确率最高仅33.20%。

Comments 9 pages, 5 figures

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(36): 29958-29966, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13844 2026-05-29 cs.CL cs.AI cs.CY cs.LG

Towards Detecting Persuasion on Social Media: From Model Development to Insights on Persuasion Strategies

检测社交媒体上的说服:从模型开发到说服策略的洞察

Elyas Meguellati, Stefano Civelli, Pietro Bernardelle, Shazia Sadiq, Irwin King, Gianluca Demartini

机构 * University of Queensland(昆士兰大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文通过开发轻量级说服文本检测模型(在SemEval 2023任务3子任务3中达到最优性能)并应用于澳大利亚联邦选举2022 Facebook广告数据集,揭示了政治竞选在不同资金策略、词汇选择、人口统计定位和选举临近时说服强度时间变化中的模式。

Journal ref Proceedings of the International AAAI Conference on Web and Social Media 20(1) (2026) 1587-1608

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15451 2026-05-29 cs.LG cs.CR stat.ME

Certified Causal Defense with Generalizable Robustness

具有泛化鲁棒性的认证因果防御

Yiran Qiao, Yu Yin, Chen Chen, Jing Ma

机构 * Case Wester Reserve University(凯斯西储大学) University of Virginia(弗吉尼亚大学)

AI总结 提出GLEAN框架,通过可认证因果因子学习解耦因果关系与虚假相关性,并设计因果认证防御策略,实现跨分布偏移域的鲁棒性泛化。

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27920 2026-05-28 cs.CV

Rethinking Video-Language Model from the Language Input Perspective

从语言输入角度重新思考视频-语言模型

Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu, Daizong Liu

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院) Huazhong University of Science and Technology(华中科技大学) Wuhan University(武汉大学)

AI总结 本文从语言输入角度出发,提出一种即插即用的框架,通过生成正负文本、属性文本推理和自加权损失,提升视频-语言模型的性能。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27894 2026-05-28 cs.CV

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

面向不完整多模态输入的统一视觉-语言模型

Xiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang, Daizong Liu, Siyi Wang, Wei Ji

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院) Guangzhou University(广州大学) Wuhan University(武汉大学) Nanjing University(南京大学)

AI总结 针对视频-语言模型在传感器失效导致模态不完整数据下的训练-测试不一致问题,提出首个统一的不完整视频-语言模型作为即插即用模块,提升多模态任务性能。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27823 2026-05-28 cs.CR cs.AI cs.CV

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

解耦对抗性提示:基于语义图的鲁棒大语言模型安全防御

Xiang Fang, Wanlong Fang

机构 * Xiang Fang(1. 方翔) Wanlong Fang(2. 方万龙)

AI总结 提出对抗性提示解耦(APD)框架,通过互信息语义分解、图谱分析和轻量级分类器,在输入处理前识别并中和恶意组件,将有害输出减少85%以上。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12955 2026-05-28 cs.AI

Text2Model: Modeling Copilots for Text-to-Model Translation

Text2Model: 用于文本到模型翻译的建模副驾驶

Serdar Kadioglu, Karthik Uppuluri, Akash Singirikonda

机构 * AI Center of Excellence, Fidelity Investments(富达投资人工智能卓越中心) Department of Computer Science, Brown University(布朗大学计算机科学系)

AI总结 本文提出Text2Model和Text2Zinc,通过统一架构和数据集、求解器无关的方式,利用多种LLM策略实现文本到组合优化与满足问题的模型翻译,并开源副驾驶和排行榜以缩小性能差距。

Comments AAAI'25 Bridge Program on Machine Learning and Operations Research CPAIOR'26 Master Class on LLMs for CP/OR

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.03808 2026-05-28 cs.LG eess.SP stat.ML

ECGadv: Generating Adversarial Electrocardiogram to Misguide Arrhythmia Classification System

ECGadv: 生成对抗性心电图以误导心律失常分类系统

Huangxun Chen, Chenyu Huang, Qianyi Huang, Qian Zhang, Wei Wang

机构 * The Hong Kong University of Science and Technology(香港科技大学) Southern University of Science and Technology, Peng Cheng Laboratory(南方科技大学鹏城实验室) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文针对基于深度神经网络的心电图诊断系统,分析心电图特性并设计两种攻击模型下的对抗攻击方案,揭示系统盲点,呼吁采取对策。

Comments Accepted by AAAI 2020

Journal ref Proceedings of the AAAI conference on artificial intelligence 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26501 2026-05-27 cs.CV cs.AI

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

揭示视觉-语言模型的脆弱性:通过纹理约束扰动和跨模态优化的多模态对抗协同

Xiang Fang, Wanlong Fang, Changshuo Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院)

AI总结 提出多模态对抗协同框架,通过纹理约束的通用对抗扰动和可学习的文本提示扰动,在黑盒设置下联合优化,揭示视觉-语言模型在多模态攻击下的脆弱性。

Comments Publish in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14640 2026-05-27 cs.CL cs.AI

Fact4ac at the Financial Misinformation Detection Challenge Task: Reference-Free Financial Misinformation Detection via Fine-Tuning and Few-Shot Prompting of Large Language Models

Fact4ac在金融虚假信息检测挑战赛中的方法:通过微调和少样本提示的大语言模型实现无参考金融虚假信息检测

Cuong Hoang, Le-Minh Nguyen

机构 * KaiNKaiho

AI总结 本文提出一种结合零样本/少样本提示和LoRA参数高效微调的大语言模型框架,用于无外部证据的金融虚假信息检测,在公开和私有测试集上分别达到95.4%和96.3%的准确率,获得竞赛第一名。

Journal ref Proceedings of the 2nd Workshop on Misinformation Detection in the Era of LLMs (MisD 2026), 20th International AAAI Conference on Web and Social Media

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05899 2026-05-27 cs.AI

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

TowerMind: 一个用于LLM作为智能体的塔防游戏学习环境与基准

Dawei Wang, Chengming Zhou, Di Zhao, Xinyuan Liu, Marci Chi Ma, Gary Ushaw, Richard Davison

机构 * Newcastle University(新castle大学) University of Auckland(奥克兰大学)

AI总结 本文提出TowerMind,一个基于塔防子类型的轻量级、多模态游戏环境,用于评估大语言模型在长期规划和决策中的能力,并揭示其与人类专家的性能差距及关键局限性。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20787 2026-05-26 cs.CV

Findings of the Counter Turing Test: AI-Generated Image Detection

反图灵测试结果:AI生成图像检测

Rajarshi Roy, Nasrin Imanpour, Ashhar Aziz, Shashwat Bajpai, Gurpreet Singh, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Amitava Das, Amit Sheth, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, Aman Chadha

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) University of South Carolina(南卡罗来纳大学) IIIT Delhi(德里IIIT) BITS Pilani Hyderabad Campus(比斯潘尼 Hyderabad 分校) IIIT Guwahati(果阿瓦提IIIT) NIT Silchar(西里char 工科院) San José State University(桑乔斯州立大学) UCLA(加州大学洛杉矶分校) Washington State University(华盛顿州立大学) Vishwakarma Institute of Information Technology(维斯瓦克arma 信息科技学院) Meta AI Amazon AI(亚马逊AI) BITS Pilani Goa(比斯潘尼 Goa 分校)

AI总结 本文通过Defactify 4.0工作坊的反图灵测试竞赛,评估了多种检测方法在区分AI生成图像与真实图像及识别具体生成模型上的性能,发现检测准确率较高但模型识别仍具挑战。

Comments Defactify4 @AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20761 2026-05-26 cs.CL

Findings of the Counter Turing Test: AI-Generated Text Detection

反图灵测试的发现:AI生成文本检测

Rajarshi Roy, Gurpreet Singh, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Amitava Das, Amit Sheth, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, Aman Chadha

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) IIIT Delhi(德里IIIT) BITS Pilani Hyderabad Campus(比斯汉学院海得拉巴校区) AI Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) IIIT Guwahati(古瓦哈提IIIT) NIT Silchar(西里char理工学院) San José State University(圣何塞州立大学) UCLA(加州大学洛杉矶分校) Washington State University(华盛顿州立大学) Vishwakarma Institute of Information Technology(维斯瓦卡马信息科技学院) Meta AI Amazon AI(亚马逊人工智能) BITS Pilani Goa(比斯汉学院果阿)

AI总结 本文通过反图灵测试(CT2)共享任务,评估了AI生成文本检测技术的有效性,发现二分类任务表现优异(F1=1.0000),但模型归因任务更具挑战性(最佳F1=0.9531),并分析了微调Transformer、集成学习等方法的优劣。

Comments Defactify4 @AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22874 2026-05-26 cs.CL

A Comprehensive Dataset for Human vs. AI Generated Text Detection

人类与AI生成文本检测的综合数据集

Rajarshi Roy, Gurpreet Singh, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Gaytri Jena, Amitava Das, Amit Sheth, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, Aman Chadha

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) IIIT Guwahati(古瓦哈提理工学院) IIIT Delhi(德里理工学院) BITS Pilani Hyderabad Campus(比什帕利 Hyderabad 分校) University of South Carolina(南卡罗来纳大学) NIT Silchar(西里 char 工程学院) San José State University(桑乔斯州立大学) UCLA(加州大学洛杉矶分校) Washington State University(华盛顿州立大学) Vishwakarma Institute of Information Technology(维斯瓦卡arma 信息科技学院) Gandhi Institute for Technological Advancement(甘地技术进步研究所) BITS Pilani Goa(比什帕利 Goa 分校) Meta AI Amazon AI(亚马逊AI)

AI总结 本文提出了一个包含73,193个文本样本的综合数据集,结合真实纽约时报文章与多个先进LLM生成的合成文本,用于区分人类与AI生成文本及归因任务,基线准确率分别为58.35%和8.92%。

Comments Defactify4 @AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19219 2026-05-26 cs.CL cs.CR

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

大型语言模型在评估中作弊了多少?基于一次性密码本的框架下的高估基准测试

Zi Liang, Liantong Yu, Shiyu Zhang, Qingqing Ye, Haibo Hu

机构 * Tech Startups(科技初创公司)

AI总结 针对大型语言模型在公开基准测试中因数据污染或训练偏差导致评估结果虚高的问题,提出基于一次性密码本加密思想的动态评估框架ArxivRoll,包含自动生成私有测试用例的SCP模块和衡量污染与偏差比例的Rugged Scores指标,实现可重复、透明且高效的评估。

Comments This paper has been accepted by AAAI 2026. We update it for adding new evaluation results for ArxivRollBench-2025a and ArxivRollBench-2026a, with the evaluation of timly models like DeepSeekV4Pro, GPT-5.5, Claude-Opus-4.7, and so on. Source code: https://github.com/liangzid/ArxivRoll/ Online Leaderboard Website: https://arxivroll.moreoverai.com/

详情

展开后加载摘要…

URL PDF HTML 收藏