arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39412cs.DB

探索性数据任务的推荐系统

Recommendation Systems for Exploratory Data Tasks

  • University of Utah(犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

Anna Fariha

中文总结 AI 辅助

本文提出探索性数据任务(EDT)推荐系统的愿景,针对行动空间巨大导致的探索瘫痪问题,沿单/多任务与单/多行动两轴规划研究议程,旨在通过智能推荐降低探索门槛并提升效率。

中文摘要 AI 辅助

一大类以数据为中心的任务是探索性的,用户在其中迭代地引导工作流程,随着新见解的出现而细化主观目标。这些探索性数据任务(EDT)由数百万具有不同专业水平的用户执行,以理解不熟悉的数据、发现趋势并识别为关键决策提供依据的证据。然而,EDT中的一个关键挑战是每一步可能采取的行动空间巨大:用户难以在数千种连接、转换和聚合操作中做出选择,从而导致“探索瘫痪”。由于EDT工作流程相互关联,每个选择都会影响后续探索,次优选择可能导致效率低下、错失见解、确认偏差和覆盖不完整。这需要智能推荐来高效地引导用户走向最优的EDT行动。我们设想推荐作为数据系统的核心能力,主动引导用户走向有前景的行动,从而降低探索性数据任务的门槛。EDT推荐具有挑战性,因为行动空间是组合性的且行动依赖于数据,这需要昂贵的物化。此外,推荐通常涉及跨相互依赖任务的行动捆绑或序列,需要跨任务协调。在本文中,我们提出了EDT推荐系统的愿景,沿两个轴展开:单任务与多任务设置,以及单行动与多行动推荐。我们概述了一个研究议程,从推荐单个EDT行动逐步发展到受约束的捆绑和行动序列,最终实现跨相互关联EDT的协调推荐。我们确定了纳入各种上下文(用户、数据、任务和生态系统)、解决效率挑战以及跨任务协调的研究方向。

英文摘要

A large class of data-centric tasks is exploratory, where users iteratively steer workflows, refining subjective goals as new insights emerge. These Exploratory Data Tasks (EDTs) are performed by millions of users with varying levels of expertise to understand unfamiliar data, discover trends, and identify evidence that informs critical decision-making. However, a key challenge in EDTs is the enormous space of possible actions that one can take at each step: users struggle to choose among thousands of joins, transformations, and aggregations, causing "exploration paralysis". Because EDT workflows are interconnected, each choice impacts subsequent exploration, and suboptimal choices can lead to inefficiency, missed insights, confirmation bias, and incomplete coverage. This calls for intelligent recommendations that efficiently guide users toward optimal EDT actions. We envision recommendation as a core capability of data systems, proactively guiding users toward promising actions and thereby lowering the barrier to exploratory data tasks. EDT recommendation is challenging because the action space is combinatorial and actions are data-dependent, which require costly materialization. Moreover, recommendation often involves bundles or sequences of actions across interdependent tasks, requiring coordination across tasks. In this paper, we present our vision of EDT recommendation systems along two axes: single-task vs. multi-task settings and single-action vs. multi-action recommendations. We outline a research agenda that progresses from recommending individual EDT actions to constrained bundles and sequences of actions, and ultimately to coordinated recommendations across interconnected EDTs. We identify research directions for incorporating various contexts (user, data, task, and ecosystem), addressing efficiency challenges, and coordinating across tasks.

↑