arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越团:比较团复形与Dowker复形在学习分析中共现数据上的应用

Beyond the Clique: Comparing Clique and Dowker Complexes for Co-occurrence Data in Learning Analytics

Koichi Yasutake, Wakana Tsuji, Hitoshi Inoue

arXiv 2610.01305首次发表:更新:

发表机构

Graduate School of Humanities and Social Sciences, Hiroshima University; Ontario Institute for Studies in Education, University of Toronto; Faculty of Business, Marketing and Distribution, Nakamura Gakuen University(广岛大学人文社会学研究科; 多伦多大学安大略教育研究所; 中村学园大学商业、营销与分销学部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文比较团复形与Dowker复形在学习分析共现数据上的表现,证明Dowker复形能保留实际共现信息、检测空洞,而团复形会消除大量真实结构,并建议根据数据类型选择Dowker复形。

AI 中文摘要

学习分析越来越倾向于将共现数据中的关系——如一段话语中的编码、一个线程中的参与者、一篇帖子上的标签——表示为单纯复形,并使用持续同调进行分析。标准方法是使用团复形。由于团复形的单纯形仅由成对网络决定,它无法区分“三个元素共同出现”与“三对元素分别出现”这两种情况。相比之下,Dowker复形将实际观察到的共现实体集合作为单纯形。我们比较了基于这两种复形的分析。首先,我们证明,在相同的1-骨架(即图结构)上,Dowker复形是团复形的子复形,由包含映射诱导的第一同调群映射是满射的,其核由从未共同出现的幻影三角形生成;也就是说,团构造只能消除空洞。然后,我们在真实数据上进行了测试。在四个Stack Exchange数据集上,46%至91%的团三角形是幻影的,并且有522至1,970个空洞被消除。在学习分析R包tna/Nestimate的示例数据上,团复形的第一贝蒂数β1在每个阈值下均为0,而Dowker复形则检测到了空洞。与保持度数的零模型相比,Dowker复形在六个数据集中的五个上表现出显著差异。此外,这种差异并未出现在实践中使用的固定阈值分析中。我们得出结论:构造方法应遵循数据类型,对于观察到的群体,Dowker复形是合适的选择;它填补了当前工具中的空白。代码可在以下网址获取:https URL。

英文摘要

Learning analytics increasingly represents the relations in co-occurrence data (codes in a window of discourse, participants in a thread, tags on a post) as simplicial complexes and analyses them with persistent homology. The standard approach uses the clique complex. Because its simplices are determined by the pairwise network alone, it cannot distinguish "three elements co-occurred together" from "each of the three pairs co-occurred separately". The Dowker complex, in contrast, takes as simplices the sets of entities actually observed to co-occur. We compare the two. First, we show that, on the same 1-skeleton, the Dowker complex is a subcomplex of the clique complex, the map on first homology induced by the inclusion is surjective, and its kernel is generated by phantom triangles that never co-occurred; that is, the clique construction can only erase holes. We also give a criterion for agreement that can be checked on triples of groups. We then test this on real data. On four Stack Exchange data sets, 46-91% of clique triangles are phantom and 522-1,970 holes are erased. On the example data of the learning-analytics R packages tna/Nestimate, the first Betti number $β_1$ of the clique complex is 0 at every threshold, whereas the Dowker complex detects holes. Against a degree-preserving null model, the Dowker complex departs strongly in five of the six data sets. This difference is invisible at the fixed thresholds used in practice. We conclude that the construction should follow the data type and that, for observed groups, the Dowker complex is the appropriate choice. Code is available at https://github.com/igu-lab/beyond-the-clique.

Commentsv3: corrects an error in the ordering of events for the human_long data; results and statements on human_long revised; code archived at doi:10.5281/zenodo.23147037

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑