arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TAHB:面向文本属性超图学习的综合基准

TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning

David Yoon Suk Kang, JungHyun Kim, Juhyun Jeon, Sang-Wook Kim

arXiv 2608.15055首次发表:更新:

发表机构

Chungbuk National University; Hanyang University(忠北国立大学; 汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出首个整合超图结构与文本属性的公开基准TAHB,含10个跨领域数据集,证实LLM增强文本语义可提升超图学习性能,为交叉领域研究提供基础。

AI 中文摘要

超图可有效建模超越成对交互的高阶组关系,而预训练语言模型(PLMs)与大语言模型(LLMs)能从文本属性中提供丰富的语义理解。然而,由于缺乏公开的文本属性超图基准,将语言模型与超图学习相结合的研究仍较为有限。为解决这一局限,我们提出TAHB(Text-Attributed Hypergraph Benchmark),即首个整合超图结构与原始文本属性的公开基准。TAHB包含来自电子商务、学术界、电影和政治网络四个领域的10个真实世界数据集,可实现对感知文本的超图表示学习的系统评估。实验结果表明,TAHB保留了真实世界超图的关键结构特性,并能一致复现现有基准中观测到的性能趋势。此外,在LLM作为增强器和LLM作为预测器的两种设置下的实验显示,LLM增强的文本语义可提升超图学习性能,而结构与文本信息的结合能为基于LLM的预测提供最佳设置。我们的基准为超图学习与语言模型交叉领域的未来研究提供了基础。

英文摘要

Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs) provide rich semantic understanding from textual attributes. However, research on combining language models with hypergraph learning remains limited due to the lack of public text-attributed hypergraph benchmarks. To address this limitation, we present TAHB (Text-Attributed Hypergraph Benchmark), the first public benchmark integrating hypergraph structures and raw textual attributes. TAHB contains 10 real-world datasets from four domains - e-commerce, academia, movies, and politics networks - enabling systematic evaluation of text-aware hypergraph representation learning. Experimental results show that TAHB preserves key structural properties of real-world hypergraphs and consistently reproduces performance tendencies observed in existing benchmarks. Furthermore, experiments under both LLM-as-Enhancer and LLM-as-Predictor settings demonstrate that LLM-enhanced textual semantics improve hypergraph learning performance, while structural and textual information jointly provide the best setting for LLM-based prediction. Our benchmark provides a foundation for future research at the intersection of hypergraph learning and language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑