TrajectoryDB:一种用于智能体轨迹的新型数据库
TrajectoryDB: A New Database for Agent Trajectories
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对AI智能体执行轨迹数据分散且缺乏专用管理的问题,本文提出轨迹应作为独立数据类型,并设计轨迹原生数据库TrajectoryDB,协同优化摄取、存储与查询处理以高效管理分析轨迹。
AI中文摘要:
AI智能体生成丰富的执行轨迹,这些轨迹捕获了它们与大型语言模型、工具和外部环境的交互。这些轨迹对于下游任务(如记忆提取、模型微调、运行时优化以及安全和成本监控)越来越有价值。然而,当前的轨迹数据分散在文件、数据库和可观测性系统中,没有围绕其独特结构和访问模式设计的持久数据管理系统。我们认为轨迹应被视为一种独特的数据类型。一条轨迹结合了层次化的执行结构、大量通常需要语义推理才能分析的文本,以及事件、中间状态和派生产物之间丰富的依赖关系和谱系。这些属性在整个数据生命周期中引入了新的需求。数据摄取必须重建并保留执行结构和谱系;存储必须高效组织庞大但高度冗余的上下文,同时维护记录之间的关系;查询处理必须联合推理结构、时间顺序、语义和谱系。因此,我们设想TrajectoryDB,一个轨迹原生的数据管理系统,它协同设计摄取、存储和查询处理,以高效管理和分析智能体执行轨迹。
英文摘要:
AI agents generate rich execution trajectories that capture their interactions with large language models, tools, and external environments. These trajectories are increasingly valuable for downstream tasks such as memory extraction, model fine-tuning, runtime optimization, and security and cost monitoring. Yet trajectory data today is fragmented across files, databases, and observability systems, with no persistent data management system designed around its unique structure and access patterns. We argue that trajectories should be treated as a distinct data type. A trajectory combines hierarchical execution structure, large volumes of text whose analysis often requires semantic reasoning, and rich dependencies and lineage among events, intermediate states, and derived artifacts. These properties introduce new requirements throughout the data lifecycle. Ingestion must reconstruct and preserve execution structure and lineage; storage must efficiently organize large but highly redundant contexts while maintaining relationships among records; and query processing must jointly reason over structure, temporal order, semantics, and lineage. We therefore envision TrajectoryDB, a trajectory-native data management system that co-designs ingestion, storage, and query processing to efficiently manage and analyze agent execution trajectories.