arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM SIGKDD Conference on Knowledge Discovery and Data Mining · 会议 · Data Mining

2026-05-29 至 2026-05-29 共收录 9
2605.30247 2026-05-29 cs.LG cs.MM

OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

OOD-GraphLLM:面向分布外泛化的药物协同预测图大语言模型

Xin Wang, Linxin Xiao, Yang Yao, Wenwu Zhu

机构 * DCST, BNRist, Tsinghua University(国防科技大学、北京理工大学、清华大学) DCST, Tsinghua University(国防科技大学、清华大学)

AI总结 针对药物协同预测中因新化合物导致的分布外偏移问题,提出OOD-GraphLLM框架,通过联合优化分子图表示与生物医学语义语言表示实现准确预测。

Comments 12 pages, 9 figures, ACM KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30027 2026-05-29 cs.CV cs.IR

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

DocRetriever:面向多模态文档检索的即插即用框架与综合基准

Ruofan Hu, Menghui Zhu, Jieming Zhu, Bo Chen, Shengyang Xu, Minjie Hong, Xiaoda Yang, Sashuai Zhou, Li Tang, Tao Jin, Zhou Zhao

机构 * Zhejiang University(浙江大学) Huawei Technologies Co., Ltd(华为技术有限公司)

AI总结 提出DocRetriever即插即用框架,通过布局感知的稀疏嵌入和推理增强的重排序器解决多模态文档检索中语义模糊和泛化瓶颈问题,并构建MultiDocR基准实现更严格评估。

Comments Accepted at KDD 2026 Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16077 2026-05-29 cs.PL cs.LG

CompilerDream: Learning a Compiler World Model for General Code Optimization

CompilerDream: 学习编译器世界模型以实现通用代码优化

Chaoyi Deng, Jialong Wu, Ningya Feng, Jianmin Wang, Mingsheng Long

机构 * School of Software, BNRist Tsinghua University Beijing China(软件学院、北师大清华大学北京中国) Tsinghua University(清华大学)

AI总结 提出基于模型的强化学习方法CompilerDream,通过编译器世界模型模拟优化pass属性并训练智能体,实现跨应用场景和语言的通用代码优化,在零样本泛化上超越LLVM内置优化。

Comments KDD 2025 camera-ready version with extended appendix. Code is available at https://github.com/thuml/CompilerDream. This update additionally fixes an issue in Table 6 where the dataset names in three rows were ordered incorrectly

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26193 2026-05-29 cs.LG cs.AI

Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection

桥接分类与重建:协同时间序列异常检测

Qideng Tang, Dai Chaofan, Wubin Ma, Yahui Wu, Haohao Zhou, Tao Zhang, Huan Li, Dalin Zhang

机构 * National Key Laboratory of Information Systems Engineering, National University of Defense Technology(信息系统工程国家重点实验室,国防科技大学) College of Systems Engineering, National University of Defense Technology(系统工程学院,国防科技大学) Zhejiang University(浙江大学) Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(空间信息感知与传输浙江大学重点实验室,杭州电子科技大学)

AI总结 提出CoAD框架,通过分类模块生成概率软掩码指导重建模块,协同利用分类与重建范式的互补优势,有效检测细微复杂异常,并在基准数据集上显著优于现有方法。

Comments 15 pages, submitted to KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22045 2026-05-29 cs.CL

DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

DLT-Corpus:面向分布式账本技术领域的大规模文本集合

Walter Hernandez Cruz, Peter Devine, Nikhil Vadgama, Paolo Tasca, Jiahua Xu

机构 * Centre for Blockchain Technologies, University College London(区块链技术中心,伦敦大学学院) School of Informatics, University of Edinburgh(信息学院,爱丁堡大学) Exponential Science Foundation(指数科学基金会)

AI总结 本文构建了DLT-Corpus,一个包含29.8亿词元、覆盖科学文献、专利和社交媒体的大规模领域语料库,并基于此分析了技术涌现模式与市场创新关联,同时发布了领域预训练模型LedgerBERT、情感分析数据集等资源。

Comments Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07044 2026-05-29 cs.CV cs.AI

PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging

PipeMFL-240K:管道磁通量泄漏成像中目标检测的大规模数据集与基准

Tianyi Qu, Songxiao Yang, Haolin Wang, Huadong Song, Xiaoting Guo, Wenguang Hu, Guanlin Liu, Honghe Chen, Yafei Ou

机构 * SINOMACH Sensing Technology \ ., Ltd Shenyang Liaoning China Institute of Science Tokyo Tokyo Japan Hokkaido University Sapporo Hokkaido Japan SINOMACH Sensing Technology \ ., Ltd Institute of Science Tokyo Hokkaido University

AI总结 为解决管道磁通量泄漏检测中缺乏大规模公开数据集和基准的问题,构建了包含249,320张图像和200,020个边界框标注的PipeMFL-240K数据集,并评估了现有目标检测器,揭示了其在长尾分布、小目标和类内变异等挑战下的性能不足。

Comments Accepted by ACM KDD 2026 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06721 2026-05-29 cs.DB

E2E: Efficient Filtered AKNN Search via Adaptive Termination

E2E:通过自适应终止实现高效的过滤AKNN搜索

Wenxuan Xia, Mingyu Yang, Wentao Li, Wei Wang

AI总结 针对带属性约束的近似k近邻搜索,提出基于早期探测阶段信息的轻量级模型实现自适应终止,在保证95%召回率下获得1.1-3.7倍加速。

Comments Accepted at KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20904 2026-05-29 cs.IR

FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets

FORGE:面向工业数据集中生成式检索的语义标识符构建

Kairui Fu, Tao Zhang, Shuwen Xiao, Ziyang Wang, Xinming Zhang, Chenchi Zhang, Yuliang Yan, Junjun Zheng, Xiangheng Kong, Shengyu Zhang, Kun Kuang, Yuning Jiang

AI总结 提出FORGE基准,通过多视角分类和离线实验研究语义标识符构建策略,并设计无需完整GR训练的新评估指标,在淘宝线上A/B测试中实现交易量提升0.35%。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10398 2026-05-29 cs.CE cs.AI

Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans

大型语言模型是否具有社会适应性?对比大型语言模型与人类的信念演化

Yu Lei, Hao Liu, Chengxing Xie, Songjia Liu, Zhiyu Yin, Canyu Chen, Guohao Li, Philip Torr, Zhen Wu

机构 * Tsinghua University(清华大学) Department of Psychological and Cognitive Sciences(心理与认知科学系) College AI(人工智能学院) School of Management(管理学院) Fudan University(复旦大学) Stevens Institute of Technology(史蒂文斯理工学院) Northwestern University(西北大学) University of Oxford(牛津大学)

AI总结 本研究提出基于社会心理学的仿真基准FairMindSim和信念-奖励对齐行为演化模型BREM,通过连续经济游戏对比人类与LLM的决策动态,发现中等能力模型表现出过度惩罚的刚性攻击性,而前沿模型随推理能力提升趋向人类式的克制与宽容。

Comments KDD 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏