发表机构
Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出从真实研究生成时间化企业场景并回放的系统,用于解决现有离线评估智能体仅基于最终静态快照的问题,可在任意时刻评估可插拔智能体。
AI 中文摘要
企业AI智能体在数据持续变化的多个应用中运行,因此答案的正确性仅取决于提问时刻存在的数据以及谁能看到这些数据。当前的离线评估仅针对单个静态快照(即事件的最终状态)评分,只能评估最终这一种情况,而事件早期的每个时刻都是不同的情境,会产生各自的现实问题和正确答案。若要将每个时刻重建为单独的快照,需为每个时刻重新配置整个租户,成本高得难以承受;且单个快照会泄露记录中隐藏的未来状态,无法呈现真实工作多应用、按时间顺序的运行方式。我们的系统同时填补了两个空白:从真实研究生成由角色驱动、随时间演化的逼真企业世界,并可在任意选定时刻回放该世界以评估任何可插拔智能体。由模式推导的时间描述驱动确定性加大语言模型(LLM)重建每条记录的过去状态;由于可查询时刻有限,所有重建内容都预计算为紧凑的差异缓存,使评估成为快速、可复现的查找过程,路径中无需模型。我们描述了该系统的设计、涵盖两个流程的架构,以及评估企业智能体的早期经验。
英文摘要
Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. Offline evaluation today grades against a single static snapshot, effectively the end of the episode. So, it can only evaluate one situation, the final one, even though every earlier moment of the episode is a different situation that invites its own realistic questions with its own correct answers. Recreating each of those moments as a separate snapshot would mean re-provisioning a whole tenant per instant, which is prohibitively costly; and even a single snapshot leaks future state hidden inside records and cannot represent the multi-app, time-ordered way real work happens. Our system closes two gaps at once: it generates a realistic, persona-driven, temporally-evolving enterprise world from real research, and replays that world at any chosen moment to evaluate any pluggable agent. A schema-inferred temporal description drives a deterministic-plus-LLM rebuild of each record's past state; because the queryable moments are finite, all rebuilds are precomputed into a compact difference cache, making evaluation a fast, reproducible lookup with no model in the path. We describe the design, an architecture spanning both flows, and early experience evaluating enterprise agents.
Comments6 pages, 3 figures, 4 tables. Accepted as a poster at SERI 2026