发表机构
Institut Polytechnique de Paris, Telecom Paris; Islamic University of Lebanon(巴黎综合理工学院,巴黎电信; 黎巴嫩伊斯兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过比较用户识别与下一域名预测任务,发现短暂浏览会话高度可识别且导航可预测,重复交互模式是主导信号,LLM语义特征增益有限。
AI 中文摘要
网页浏览通常看起来是短暂的:用户访问几个网站,完成一项任务,然后离开。然而,即使是短暂的浏览活动片段也可能包含丰富且结构化的行为信号。在这项工作中,我们对两个互补的行为推断任务进行了比较性实证研究:会话级用户识别和下一域名预测。这两个任务均源自相同的清洗后事件流,并在大规模匿名浏览轨迹上进行评估,会话划分和分割根据每个任务的时间要求进行了调整。对于用户识别,我们评估了基于会话级行为和域名特征的经典模型与神经模型。对于下一域名预测,我们将基于图的建模与大型语言模型(LLMs)相结合。实验结果表明,短暂的浏览会话具有高度的可识别性,而未来的导航行为则可以从长期交互结构与近期行为上下文的结合中高度预测。此外,LLM衍生的语义特征相比纯结构和序列模型仅带来边际收益,表明重复交互模式在评估的网页浏览设置中仍是主导的预测信号。这些发现强调了交互历史在浏览轨迹中对用户可识别性和导航可预测性的显著贡献程度。
英文摘要
Web browsing often appears ephemeral: users visit a few websites, complete a task, and move on. However, even short fragments of browsing activity can contain rich and structured behavioral signals. In this work, we conduct a comparative empirical study of two complementary behavioral inference tasks: session-level user identification and next-domain prediction. Both tasks are derived from the same cleaned event stream and evaluated on large-scale anonymous browsing traces, with sessionization and splitting adapted to the temporal requirements of each task. For user identification, we evaluate classical and neural models operating on session-level behavioral and domain features. For next-domain prediction, we combine graph-based modeling with Large Language Models (LLMs). Experimental results show that short browsing sessions are highly identifiable, while future navigation actions are highly predictable from long-term interaction structure combined with recent behavioral context. Furthermore, LLM-derived semantic features yield only marginal gains over purely structural and sequential models, indicating that repeated interaction patterns remain the dominant predictive signal in the evaluated web-browsing setup. These findings highlight the extent to which interaction history substantially contributes to both user identifiability and navigation predictability in browsing traces.