发表机构
Technical University of Munich; Nemetschek Group(慕尼黑工业大学; Nemetschek集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推出IFCMemoryBench基准,评估基于LLM的智能体在BIM信息检索中的长期记忆,发现现有通用记忆系统在该场景下表现差,存在领域迁移差距。
AI 中文摘要
长期记忆正成为基于大语言模型(LLM)的智能体的核心能力,但现有评估大多仅测试开放域或基于人设场景下的对话式召回。本文提出,更严格的测试是智能体能否在活跃、结构化的特定领域环境中行动时复用过往会话的信息。我们在建筑信息模型(BIM)这一专业工程工作流中研究该问题,在此场景中,智能体需查询大型IFC模型,同时依赖项目规范、客户决策及工程惯例,这些内容常于对话中讨论但未包含在模型内。我们推出IFCMemoryBench,这一用于评估基于LLM的BIM信息检索中长期记忆能力的基准。该基准包含19个项目中的143项多会话任务及4016个过往会话,源自IFC-Bench v2中的信息不完备问题。每项任务会在早期对话中植入缺失的项目上下文,后续提出的探测问题需结合记忆的上下文与活跃的IFC查询才能解答。我们的评估框架将记忆性能分解为录入、检索和利用三个环节,并通过经专家验证的LLM评判器同时测量答案质量与记忆质量。我们评估了具有代表性的基于向量、图和文件的记忆系统。在符合部署实际的录入范围内,最强的系统仅达到32.4%的答案准确率;在经预言机过滤的录入或更强的探测智能体下,其准确率仍低于60%。分析显示,当前通用记忆系统常检索到主题相关的上下文,但将项目知识存储为不完整或碎片化的事实。这些结果揭示了智能体记忆中的领域迁移差距,表明可靠的专业智能体需要能关联对话、项目知识与结构化模型实体的领域感知记忆表征。
英文摘要
Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings. We argue that a stronger test is whether an agent can reuse information from prior sessions while acting over a live, structured, domain-specific environment. We study this problem in Building Information Modelling (BIM), a professional engineering workflow where agents must query large IFC models while also relying on project specifications, client decisions, and engineering conventions often discussed in conversation but absent from the model. We introduce IFCMemoryBench, a benchmark for evaluating long-term memory in LLM-based BIM information retrieval. IFCMemoryBench contains 143 multi-session tasks across 19 projects and 4,016 prior sessions, derived from incomplete-information questions in IFC-Bench v2. Each task seeds missing project context across earlier conversations and later asks a probe question that can be answered only by combining remembered context with live IFC queries. Our evaluation framework decomposes memory performance into ingestion, retrieval, and utilization, and measures both answer quality and memory quality with expert-validated LLM judges. We evaluate representative vector-, graph-, and file-based memory systems. The strongest system achieves only 32.4% answer accuracy under a deployment-realistic ingestion scope, and remains below 60% under oracle-filtered ingestion or a stronger probe agent. Analysis shows that current general-purpose memory systems often retrieve topically relevant context but store project knowledge as incomplete or fragmented facts. These results reveal a domain-transfer gap in agent memory and suggest that reliable professional agents require domain-aware memory representations linking conversations, project knowledge, and structured model entities.
CommentsKDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI