arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

ACM SIGKDD Conference on Knowledge Discovery and Data Mining · 会议 · Data Mining

至 收录 2566
2505.22071 2026-07-14 physics.geo-ph 版本更新

Ocean-E2E: Hybrid Physics-Based and Data-Driven Global Forecasting of Extreme Marine Heatwaves with End-to-End Neural Assimilation

Ocean-E2E:基于混合物理和数据驱动的全球极端海洋热浪端到端神经同化预测

Ruiqi Shu, Ruijian Gou, Yanfei Xiang, Xiaomeng Huang

AI总结 研究聚焦全球极端海洋热浪端到端预测,创建混合框架Ocean-E2E,基于物理性质通过端到端数据同化及模拟相关作用提高预测能力,能进行端到端和区域高分辨率预测,优于现有模型,为气候极端事件预测提供框架。

Comments Accepted by KDD 2026

URL PDF HTML 收藏
2606.03631 2026-07-13 cs.LG cs.AI 版本更新

AnchorMoE: Interpretable Time Series Classification via Anchor-Routed MoE

AnchorMoE: 基于锚点路由的混合专家模型实现可解释时间序列分类

Tao Xie, Zexi Tan, Haoyi Xiao, Mengke Li, Yiqun Zhang, Yang Lu, Cuie Yang, Yiu-ming Cheung

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) School of Computer Science and Technology, Guangdong University of Technology(广东工业大学计算机科学与技术学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Informatics, Xiamen University(厦门大学信息学院) State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University(东北大学过程工业综合自动化国家重点实验室) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系)

AI总结 提出AnchorMoE框架,利用混合专家架构对局部补丁进行多视角表示并路由至专门专家,通过加性分解实现前向可解释性,并引入几何正交约束和不确定性感知门控机制提升稀疏信号下的分解可靠性与噪声抑制。

Comments Accepted by KDD 2026

URL PDF HTML 收藏
2602.10392 2026-07-13 cs.LG 版本更新

Tensor Methods: A Unified and Interpretable Approach for Material Design

张量方法:一种统一且可解释的材料设计方法

Shaan Pakala, Aldair E. Gongora, Brian Giera, Evangelos E. Papalexakis

机构 * University of California, Riverside(加州大学河滨分校) Dept. of Computer Science & Engineering(计算机科学与工程系) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Materials Engineering Division(材料工程 division) Data Science Institute(数据科学研究所)

AI总结 提出使用张量补全方法作为材料设计的统一框架,兼具可解释性和预测性能,在非均匀采样下优于传统机器学习,最高提升5%的R²并减半分布外误差。

Comments To appear in the ACM SIGKDD 2026 AI for Sciences track

Journal ref KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, 2026, pp. 11740-11749

URL PDF HTML 收藏
2607.08555 2026-07-10 cs.LG 新提交

CAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal Consistency

CAAD:通过多尺度对齐和结构因果一致性进行因果感知多变量时间序列异常检测

Xin Wang, Yunshi Wen, Yanan He, Haotian Xu, Youlan Zhao, Michel Ferreira Cardia Haddad, Tengfei Ma

机构 * Stony Brook University(纽约州立大学石溪分校) Rensselaer Polytechnic Institute(伦斯勒理工学院) Yale University(耶鲁大学) Queen Mary University of London(伦敦大学玛丽皇后学院)

AI总结 针对复杂工业系统异常检测中常忽略内部因果关系的问题,提出CAAD框架,通过外生变量持续验证格兰杰因果一致性,利用多尺度对齐和梯度矩阵监测因果关系,在真实工业数据集实验中实现高精度异常检测,优于多数基线方法。

Comments Accepted at KDD 2026 (Research Track)

URL PDF HTML 收藏
2607.08475 2026-07-10 cs.LG 新提交

Frequency-Domain Multi-Modality Transportation Modeling

频域多模态交通建模

Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, Renhe Jiang

机构 * Southern University of Science and Technology(南方科技大学) The University of Tokyo(东京大学)

AI总结 针对多模态交通预测难题,提出频域多模态建模FreMo,通过模态频域滤波器细化频谱、频率引导协同积分器聚合跨模态信息,实现自适应和选择性跨模态协同,实验证明其性能优于现有基线。

Comments Accepted by KDD 2026 Research Track

URL PDF HTML 收藏
2607.08071 2026-07-10 cs.CL cs.AI cs.LG 新提交

COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation

COBART:用于广告标题生成的可控、优化、双向和自回归变换器

Yashal Shakti Kanungo, Gyanendra Das, Pooja A, Sumit Negi

机构 * Amazon(亚马逊)

AI总结 研究针对广告标题生成难题,提出用前缀控制令牌结合BART微调的方法,可控制标题长度,适应不同格式与要求,实验显示该方法能提升Rouge-L和估计CTR,相比基线有显著提高。

Comments 10 pages, 5 figures, 5 tables. Published in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '22). This is the author's accepted version; the definitive Version of Record is available at https://doi.org/10.1145/3534678.3539069

Journal ref Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '22), August 14-18, 2022, Washington, DC, USA, pp. 3127-3136

URL PDF HTML 收藏
2607.08043 2026-07-10 cs.SE cs.AI 新提交

Aleena: Alignment Agent for Research Software Engineering Collaborations

Aleena:用于研究软件工程协作的对齐代理

Kshitij Dani, Cordero Core, Landung Setiawan, Carlos Garcia Jurado Suarez, Anshul Tambay, Vani Mandava, Anant Mittal

机构 * eScience Institute(eScience研究院) University of Washington(华盛顿大学)

AI总结 研究软件协作中决策易失依据致相关人员心智模型不同,提出智能AI可支持对齐与跟踪。介绍开源的Aleena,以GitHub为协作平台,转化交互为结构化记录,揭示风险等,还阐述其动机、设计、原型及场景。

Comments 8 pages, 5 figures. AgenticSE @ KDD '26: Agentic Software Engineering (SE 3.0): The Rise of AI Teammates, KDD 2026 Workshop

URL PDF HTML 收藏
2607.08092 2026-07-10 quant-ph cs.CV 新提交

Equivariant Quantum Clustering with Differential Privacy: Parameter-Efficient Privacy-Preserving Analysis Across Heterogeneous Sensitive Datasets

具有差分隐私的等变量子聚类:跨异构敏感数据集的参数高效隐私保护分析

B. M. Taslimul Haq, Md Arifur Rahman, Tawfiq Al Islam Foysal, Abdullah Al Noman, Abir Ahmed

AI总结 研究如何在隐私保护下对异构敏感数据集进行聚类分析,提出等变量子聚类框架EQC,结合对称感知量子电路与差分隐私,采用p4m等变参数共享降复杂度,实验表明其在多个数据集上表现良好,为隐私保护聚类提供实用框架。

Comments 24 pages, 10+ tables, multiple figures, research article. Introduces Equivariant Quantum Clustering (EQC) integrating differential privacy with parameter-efficient quantum circuits for privacy-preserving clustering. Evaluated on NSL-KDD, CERT Insider Threat v6.2, and Synthetic MIMIC-III datasets

Journal ref Journal of AI ML DL, Vol. 1, No. 1, 2025, pp. 1-24

详情
URL PDF HTML 收藏
2606.06104 2026-07-10 cs.LG 版本更新

A Sliced-Wasserstein Framework on Correlation Matrices for EEG Decoding

用于脑电图解码的相关矩阵切片Wasserstein框架

Chen Hu, Rui Wang, Jiale Zhou, Jingjun Yi, Shaocheng Jin, Yidong Song, Yefeng Zheng

机构 * Westlake University(西湖大学) School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Jiangnan University(江南大学) Sun Yat-sen University(中山大学)

AI总结 提出基于拉回欧几里得度量的切片Wasserstein框架,实例化两种相关矩阵切片Wasserstein差异,并构建脑电图解码的域泛化方法,在三个数据集上验证了分布偏移下的泛化能力提升。

Comments Accepted by KDD 2026

URL PDF HTML 收藏
2607.07504 2026-07-09 cs.AI 新提交

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

大语言模型生成的技能能否造就更好的人工智能数据科学家?跨数据科学工作流程的组件消融

Wei-Jung Huang

机构 * Independent Researcher(独立研究者)

AI总结 研究大语言模型生成的技能在数据科学工作流程中的作用,通过在四个阶段测试及组件消融实验发现,完整生成技能及消融后的技能变体与任务单独提示相比,均未显著提升性能,提醒勿将其作为默认单次提示策略。

Comments KDD 2026 Workshop on AI Data Scientist

Journal ref KDD 2026 Workshop on AI Data Scientist

URL PDF HTML 收藏
2412.12152 2026-07-09 cs.LG cs.AI

GraphTool-Instruction: Revolutionizing Graph Reasoning in LLMs through Decomposed Subtask Instruction

GraphTool-Instruction: 通过分解子任务指令革新大语言模型中的图推理

Rongzheng Wang, Shuang Liang, Qizhi Chen, Jiasheng Zhang, Ke Qin

机构 * University of Electronic Science and Technology of China(电子科技大学)

AI总结 GraphTool-Instruction通过分解图推理任务为三个子任务,提升LLMs在图推理中的性能,实现SOTA效果。

Comments 22 pages, have been accepted by KDD 2025

URL PDF HTML 收藏
2607.05613 2026-07-08 cs.LG 新提交

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

SafeImpute:通过共形选择实现可靠的临床数据插补

Xinrui He, Mengting Ai, Junting Wang, Curtiss B. Cook, Jingrui He

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Mayo Clinic(梅奥诊所)

AI总结 研究临床数据插补问题,提出SafeImpute框架,通过构建事件图、利用双关系GNN和自适应融合学习插补,并结合共形选择控制不可接受误差,实验证明该方法在插补准确性和可靠误差控制上优于基线。

Comments Accepted at KDD 2026. Author accepted manuscript

URL PDF HTML 收藏
2607.05242 2026-07-07 cs.LG cs.AI cs.IR 新提交

CanniUplift: A Holistic Framework for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

CanniUplift:缓解电商增益建模中商家蚕食与激励蚕食的整体框架

Zuwang He, Shihao Shu, Yuli Qu, Hanyu Gao, Ziliang Zhang, Diwei Chen, Xiangda Yan, Buyu Gao, Tanchao Zhu, Yumeng Li, Junxiong Zhu

机构 * Taobao & Tmall Group of Alibaba(阿里巴巴淘宝及天猫集团)

AI总结 针对传统电商增益模型违反SUTVA的两类蚕食问题,提出含PGA、RDD和Treat-Attention的统一框架,实验与线上部署均验证其可显著提升平台增量GMV与ROI。

Comments Accepted to KDD 2026, 12 pages, 4 figures

URL PDF HTML 收藏
2607.04557 2026-07-07 cs.LG cs.AI q-bio.QM 新提交

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

通过对齐患者特异性知识图谱和基因水平扰动表示来预测治疗结果

Dongmin Bang, Sugyun An, Inyoung Sung, Ilho Yun, Sun Kim, Sangseon Lee

机构 * Interdisciplinary Program in Bioinformatics, Seoul National University(首尔国立大学生物信息学跨学科项目) AIGENDRUG Co., Ltd.(爱真药物有限公司) BK21 FOUR Intelligence Computing, Seoul National University(首尔国立大学BK21四号智能计算) Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔国立大学人工智能跨学科项目) Department of Artificial Intelligence, Inha University(仁荷大学人工智能系)

AI总结 针对临床响应标签和治疗后分子图谱稀缺阻碍治疗反应预测的问题,提出PREDIKTOR框架,结合个性化网络视图与可转移转录组扰动视图预测临床药物反应,提升性能并支持精准肿瘤学。

Comments 12 pages, 5 figures, 5 tables. Accepted at BIOKDD 2026, held in conjunction with ACM SIGKDD 2026

URL PDF HTML 收藏
2607.02945 2026-07-07 cs.PF 新提交

Optimus: A Generic Operator-Level PyTorch Model Transformation Framework

Optimus:一个通用的算子级PyTorch模型转换框架

Menglu Yu, Jiaqi Xu, Yuzhen Huang, Yanbo Liang, Jia Liu, Shuai Yang, Jason Ansel, Elias Ellison, Edward Yang, Brian Hirsh, Jia Chen Ren, Will Feng, Oguz Ulgen, Xu Zhao, Daohang Shi, Huaqing Xiong, Quanyu Zhu, Mingming Ding, Junqing Zhou, Ruilin Chen, Yuhang Yang, Chi-Keung Luk

AI总结 针对大规模工业应用中深度学习模型架构复杂多样、人工优化不可行的问题,介绍基于PyTorch 2.x的通用模型转换框架Optimus,用预定义模式和贪心搜索算法,可提升模型性能。

Journal ref In Proceedings of the 2026 ACM SIGKDD Conference on Knowledge Discovery and Data Mining

URL PDF HTML 收藏
2607.02928 2026-07-07 cs.LG 新提交

CoFEND: A Cross-Modal Fusion End-to-End Network for Cold-Start Drug-Drug Interaction Prediction

CoFEND:用于冷启动药物-药物相互作用预测的跨模态融合端到端网络

Di Wu, Hongyi Sun, Haichao Xu, Jia Chen, Zhong Chen, Jie Yang

机构 * College of Computer and Information Science, Southwest University(西南大学计算机与信息科学学院) Beihang University(北京航空航天大学) School of Computing, Southern Illinois University Carbondale(南伊利诺伊大学卡本代尔分校计算机学院) School of Physics and Electronic Science, Zunyi Normal University(遵义师范学院物理与电子科学学院)

AI总结 针对新药冷启动药物-药物相互作用预测难题,提出CMF-ELN。利用多模态信息构建知识图谱,设计图自动编码器融合跨模态相似性,采用两阶段可解释性方案,提升预测准确性与机制解释性。

Comments 11 pages, 2 figures, accepted by KDD 2026

URL PDF HTML 收藏
2607.02703 2026-07-07 cs.SE cs.AI cs.DC cs.MA 新提交

LLMoxie: Exploring Agentic AI for Scientific Software Development

LLMoxie:探索用于科学软件开发的智能AI

Landung Setiawan, Anant Mittal, Cordero Core, Anshul Tambay, Carlos Garcia Jurado Suarez, David A. C. Beck, Andrew J. Connolly, Vani Mandava

机构 * eScience Institute University of Washington(eScience研究院华盛顿大学)

AI总结 研究针对科学软件开发中现有AI编码代理的不足,介绍LLMoxie平台及其插件生态系统,通过实践总结多域研究软件工程师中心采用智能AI的挑战、平台设计及操作经验,提升AI编码代理能力。

Comments 9 pages, 4 figures. Accepted to ACM SIGKDD 2026 Workshop: Agentic AI for Scientific and Societal Advances (SciSoc Agents and LLMs). Describes an agentic AI platform for scientific software engineering with governed multi-cloud inference, structured multiagent workflows, and domain-aware coding support (cs.SE, cs.MA, cs.AI)

URL PDF HTML 收藏
2605.18580 2026-07-07 cs.AI cs.LG 版本更新

When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State

当结果看似正确但纪律却失败:基于轨迹的评估在隐藏对手状态下的应用

Peiying Zhu, Sidi Chang

机构 * Blossom AI Blossom AI Labs(Blossom AI 实验室)

AI总结 本文提出了一种基于轨迹的评估方法,用于评估在隐藏对手状态下的行为纪律稳定性,通过轨迹诊断、机制分离和转移测试来改进强化学习策略,特别是在酒店定价和隐藏预算竞标任务中。

Comments Accepted to the KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI

URL PDF HTML 收藏
2604.00392 2026-07-07 cs.SE cs.AI 版本更新

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

EvolveTool-Bench:评估LLM生成的工具库作为软件 artifact 的质量

Alibek Kaliyev, Artem Maryanskyy

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Uber Technologies(优步科技公司)

AI总结 本文提出EvolveTool-Bench,通过评估LLM生成的工具库在软件工程流程中的质量,揭示任务完成率之外的软件质量风险,强调需将工具库视为首要软件 artifact。

Comments 11 pages, 4 figures; accepted at KDD 2026 Workshop on Agentic AI Evaluation and Trustworthiness

URL PDF HTML 收藏
2607.02115 2026-07-03 cs.IR 新提交

Planning over Matrix-Factorization MDPs for Candidate Generation

基于矩阵分解MDP的候选生成规划

Mikhail Trapeznikov, Maksim Utushkin

AI总结 将推荐系统的用户旅程建模为MDP,通过折叠更新用户状态进行规划,实验表明单步前瞻即可显著提升固定嵌入下的检索效果。

Comments Accepted to the 5th Workshop on End-to-End Customer Journey Optimization at KDD 2026. 6 pages, 3 figures, 2 tables

URL PDF HTML 收藏
2607.01773 2026-07-03 cs.AI 新提交

Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis

通过检索基础的形式概念分析实现可验证的知识扩展

Yujin Yang, Heejung Lee

机构 * Hanyang University(汉阳大学)

AI总结 提出一种检索增强的小语言模型框架,利用形式概念分析作为符号验证循环,通过种子属性和检索验证实现知识扩展,在罕见共济失调数据集上评估了关系F1和蕴含F1。

Comments 8 pages, 2 figures, Accepted to the 8th epiDAMIK ACM SIGKDD International Workshop on Epidemiology meets Data Mining and Knowledge Discovery (epiDAMIK 2026)

URL PDF HTML 收藏
2607.01485 2026-07-03 cs.IR 新提交

CoPersona: Collaborative Persona Graphs for Robust LLM Personalization

CoPersona: 面向鲁棒LLM个性化的协作人格图

Yangtian Zhang, Leyao Wang, Hiren Madhu, Ngoc Bui, Walter Roznyatovskiy, Rex Ying

AI总结 针对用户历史稀疏和偏差导致的个性化脆弱问题,提出CoPersona框架,通过多面人格图从行为相似用户中借调信号,结合非参数检索与参数图推理的双分支架构,在多个领域和模型规模上超越强基线。

Comments Accepted at KDD '26. 12 pages, 5 figures, 8 tables

URL PDF HTML 收藏
2602.14110 2026-07-03 cs.IR 版本更新

MixFormer: Co-Scaling Up Dense and Sequence in Industrial Recommenders

MixFormer: 在工业推荐系统中协同扩展稠密与序列

Xu Huang, Hao Zhang, Zhifang Fan, Yunwen Huang, Zhuoxing Wei, Zheng Chai, Jinan Ni, Yuchao Zheng, Qiwei Chen

AI总结 提出MixFormer统一Transformer架构,联合建模序列行为与特征交互,通过统一参数化实现稠密容量与序列长度的协同扩展,并引入用户-物品解耦策略提升工业实用性。

Comments Accepted by KDD 2026

URL PDF HTML 收藏
2607.00958 2026-07-02 cs.LG 新提交

LeNEPA: No-Augmentation Next-Latent Prediction for Time-Series Representation Learning

LeNEPA:无增强的下一潜在预测用于时间序列表示学习

Alexander Chemeris, Ming Jin, Randall Balestriero

机构 * Langotime South Africa(Langotime南非) Griffith University Australia(澳大利亚格里菲斯大学) Brown University United States(美国布朗大学)

AI总结 提出LeNEPA,一种无需数据增强的下一潜在预测架构,通过SIGReg正则化和轻量投影空间,在固定配方测试中跨数据集保持性能,优于ECG调优的JEPA。

Comments 9 pages, 4 figures, 6 tables; accepted by the 12th Mining and Learning from Time Series (KDD MILETS 2026); source code and artifacts: https://github.com/langotime/lenepa-milets-2026

URL PDF HTML 收藏
2607.00956 2026-07-02 cs.LG cs.AI 新提交

Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations

Aionoscope:时间序列表示中潜在状态可访问性的调试

Alexander Chemeris, Ming Jin, Randall Balestriero

机构 * Langotime South Africa(Langotime 南非) Griffith University(格里菲斯大学) Brown University(布朗大学)

AI总结 提出Aionoscope,一种基于生成器的诊断工具,用于调试冻结时间序列表示中的潜在状态可访问性,揭示粗粒度和细粒度可访问性之间的不匹配。

Comments 9 pages, 4 figures. Accepted by the 12th Mining and Learning from Time Series (KDD MILETS 2026). Interactive results: https://aionoscope.langotime.ai/ . Source artifacts: https://github.com/langotime/aionoscope/ and https://github.com/langotime/aionoscope-benchmarks/

URL PDF HTML 收藏
2607.00280 2026-07-02 cs.LG cs.CY econ.EM stat.AP 新提交

Understanding Guest Preferences and Optimizing Two-sided Marketplaces: Airbnb as an Example

理解客人偏好与优化双边市场:以Airbnb为例

Yufei Wu, Daniel Schmierer

机构 * Airbnb, Inc.(Airbnb公司)

AI总结 结合经济建模与因果推断,分析客人对价格等因素的响应及偏好异质性,以优化定价工具和个性化推荐,提升市场匹配效率。

Comments 5 pages, 3 figures. Presented at the KDD 2024 Workshop on Two-Sided Marketplace Optimization, Barcelona, Spain

URL PDF HTML 收藏
2607.00398 2026-07-02 cond-mat.str-el cs.AI cs.ET 新提交

Holographic Quantum Transformer: A Generalist Neuro-Symbolic Architecture for Solving Frustrated Systems via Generative Attention

全息量子变换器:一种通过生成式注意力解决受挫系统的通用神经符号架构

Xingran Guo, Tiaojie Xiao, Jie Liu, Keqin Li

机构 * National University of Defense Technology(国防科技大学) State University of New York at New Paltz(纽约州立大学新帕尔茨分校)

AI总结 提出全息量子变换器(HQT),利用全局自注意力解决非局域纠缠模式,在受挫J1-J2海森堡模型上达到高精度,并实现零样本尺寸外推协议。

Comments 10 pages, accepted to KDD '26

Journal ref In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD '26), August 09-13, 2026, Jeju Island, Republic of Korea

URL PDF HTML 收藏
2606.03137 2026-07-02 cs.AI 版本更新

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

Think-Before-Speak: 从内部评估到多智能体社会模拟中的公开表达

Kaiqi Yang, Tai-Quan Peng, Sanguk Lee, Hui Liu

机构 * Michigan State University(密歇根州立大学) Hankuk University of Foreign Studies(韩国民法大学)

AI总结 提出TBS框架,通过分离智能体的内部推理与公开话语生成,模拟从内部评估到公开表达的路径,并在气候政策讨论中验证其机制敏感性。

Comments 8 pages of main content, 14 pages including references and appendices, 3 figures. Accepted to the KDD'26 Workshop on SciSoc Agents & LLMs

URL PDF HTML 收藏
2606.29180 2026-07-01 cs.AI 版本更新

Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings

知识图谱中的图到图语义相似度测量:知识图谱嵌入的实证评估

Seungryeol Baek, Wooseok Sim, Hogun Park

机构 * Sungkyunkwan University(成均馆大学)

AI总结 本文通过构建语义匹配数据集,比较基于文本、结构和知识图谱嵌入的方法,提出EmbPairSim和AvgEmbSim评分函数,实验表明EmbPairSim在参数更少的情况下优于Sentence-BERT。

Comments 9 pages, 2 figures, 6 tables. Accepted as a poster at The 2nd Frontiers in Graph Machine Learning for the Large Model Era (GMLLM'26) Workshop, co-located with KDD 2026

URL PDF HTML 收藏
2606.30992 2026-07-01 stat.ME cs.LG stat.AP 新提交

Hierarchical Clustering As a Novel Solution to the Notorious Multicollinearity Problem in Observational Causal Inference

层次聚类作为观测因果推断中多重共线性问题的新解决方案

Yufei Wu, Zhiying Gu, Alex Deng, Jacob Zhu, Linsha Chen

机构 * Airbnb, Inc.(Airbnb公司)

AI总结 针对观测因果推断中多重共线性导致无法分离变量影响的问题,提出基于层次聚类聚合数据以缓解共线性的方法,并通过营销应用验证其有效性。

Comments Presented at the KDD 2023 Workshop on Causal Inference and Machine Learning in Practice, Long Beach, CA; also presented at the 2023 Joint Statistical Meetings

URL PDF HTML 收藏