RunningTab: 通过环境侧标签实现直接工作区交互
RunningTab: Direct Workspace Interaction with Environment-Side Tabs
浏览论文内容
中文总结 AI 辅助
RunningTab通过环境侧标签跟踪任务需求与文件读取状态,解决LLM智能体在直接工作区交互中遗漏关键信息的问题,显著提升交付物完整性。
中文摘要 AI 辅助
许多知识工作从工作区已有的文件中产生新的交付物,而LLM智能体正开始接管此类工作。通过直接语料库交互,智能体可以从终端搜索和读取任何文件,无需索引,以这种方式从多个文件中生成交付物,我们称之为直接工作区交互(DWI)。然而,访问文件只是任务的一半:没有任何机制跟踪任务要求什么、已读取什么、以及列出但未打开的文件,所有这些都会从上下文窗口中滑过而不留痕迹,因此智能体可能提取了一个图表,却仍然交付了缺少该图表的报告。为解决此问题,我们提出了RunningTab,一个框架,它为直接工作区交互配备了一个环境侧标签:一个按任务记录任务仍欠缺内容的记录,由环境与智能体一同维护。具体而言,智能体添加其需求,而环境将每次文件读取记录为带来源的摘录,并将每个列出但未打开的文件记录为候选;智能体随后可以查看每个需求及其最佳匹配的摘录和未打开的候选,根据匹配内容解决需求,或附上理由将其搁置,并且如果它试图在仍有需求未解决时完成,它会在完成检查中收到这些需求。我们在三个基准测试上使用三个LLM验证了RunningTab,它始终优于纯DWI和将记录保存在模型中的基线,而其标签通常包含交付物所需的数值,一旦看到这些数值。
英文摘要
Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。