arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03141cs.DB

当模型“吞噬”栈时会发生什么?重新思考数据智能体应对“苦涩教训”的研究议程

What Happens When the Model Eats the Stack? Rethinking the Research Agenda for Data Agents to Withstand the Bitter Lesson

Liana Patel, Siddharth Jha, Negar Arabzadeh, Carlos Guestrin, Ion Stoica, Matei Zaharia

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型内化数据智能体能力的“苦涩教训”问题,该文提出应聚焦持久语义上下文相关研究,以支撑数据智能体在庞大知识语料库上的运行。

中文摘要 AI 辅助

“苦涩教训”为数据系统领域提出了一个根本性问题:端到端训练的大语言模型(LLMs)正快速内化此前需精心设计数据智能体才能实现的新能力。基于实证见解,我们认为随着模型持续改进,许多为弥补模型在特定任务上局限性而设计的系统层将越来越多地被模型本身所涵盖。我们确定了持久的研究机会,这些机会在于通过关于数据环境的精选上下文信息(我们称之为持久语义上下文)为跨多个查询的数据智能体提供支持。我们发现这些上下文层在提升数据智能体性能方面展现出巨大潜力,但也带来了重大系统挑战。因此,未来数据系统的一个关键需求在于将持久语义上下文作为一等抽象原生提供服务,以支持能在庞大复杂知识语料库上运行的高性能数据智能体。为实现这一愿景,我们概述了令人兴奋的新研究机会,包括设计高效的上下文数据结构、存储方法、压缩技术以及语义一致性协议,以确保所存储上下文知识的完整性与正确性。

英文摘要

The bitter lesson poses an existential question for the data systems community, whereby large language models (LLMs) trained end-to-end are rapidly internalizing new capabilities that previously required carefully engineered data agents. Guided by empirical insights, we argue that as models continue to improve, many proposed system layers designed to compensate for model limitations on a given task will increasingly be subsumed by the model itself. We instead identify enduring research opportunities, which lie in supporting data agents across many queries with curated contextual information about the data environment, which we call persistent semantic context. We find that these context layers demonstrate strong promise for improving data agent performance, but they also raise significant system challenges. Thus, a key requirement for future data systems will lie in natively serving persistent semantic contexts as a first-class abstraction in order to enable capable data agents working over huge, complex knowledge corpora. Towards this vision, we outline exciting new research opportunities, including designing efficient context data structures, storage methods, compression techniques, and semantic consistency protocols, to ensure integrity and correctness of the stored contextual knowledge.

发表机构

  • UC Berkeley(加州大学伯克利分校)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑