从场景图到答案:面向自动驾驶的选择性神经符号推理
From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving
浏览论文内容
中文总结 AI 辅助
提出一种查询自适应的神经符号推理框架,通过时空场景图和符号执行器优先处理精确查询,仅在必要时调用LLM,在NuScenes-QA上提升准确率与效率。
中文摘要 AI 辅助
自动驾驶问答需要对结构化场景信息进行推理,然而现有的视觉-语言方法大多将异构推理操作委托给单一的神经推理过程。我们认为这种统一策略忽略了一个根本性区别:某些查询允许精确的符号解,而其他查询则需要语义解释。我们引入了一个查询自适应的神经符号推理框架,该框架根据查询的性质显式分配计算。其核心是一个层次化的时空场景图(STSG),它将持久的对象身份与帧特定状态分离,并将空间关系和时序转换表示为显式的有向结构。给定一个查询,符号执行器首先尝试通过精确的图操作来解决它;只有当符号执行弃权(不执行)时,才会调用大语言模型(LLM)进行语义推理。对于这些未解决的查询,查询条件的图检索和证据过滤保留了关系方向、时间局部性和对象语义,为LLM提供紧凑且经过验证的任务相关证据。这种设计将LLM的角色从通用推理引擎转变为有针对性的语义推理器,同时允许确定性计算被精确且高效地处理。我们在oracle感知设置下,对nuScenes v1.0-mini所有十个场景中的5,916个NuScenes-QA问题评估了该框架。完整系统使用GPT-5.4-mini实现了80.63%的总体准确率,比相应的仅LLM配置提高了5.48个百分点;使用DeepSeek-V4-Flash时,改进达到6.64个百分点。最大的增益出现在计数问题上,分别提高了10.20和12.61个百分点。这些结果表明,选择性推理提高了准确性和推理效率。
英文摘要
Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some queries admit exact symbolic solutions, while others require semantic interpretation. We introduce a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query. At its core is a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures. Given a query, a symbolic executor first attempts to resolve it through exact graph operations; only when symbolic execution abstains is an LLM invoked for semantic reasoning. For these unresolved queries, query-conditioned graph retrieval and evidence filtering preserve relation direction, temporal locality, and object semantics, providing the LLM with compact and verified task-relevant evidence. This design shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently. We evaluate the framework on 5,916 NuScenes-QA questions across all ten scenes of nuScenes v1.0-mini under an oracle-perception setting. The complete system achieves 80.63 percent overall accuracy with GPT-5.4-mini, improving over the corresponding LLM-only configuration by 5.48 percentage points; with DeepSeek-V4-Flash, the improvement reaches 6.64 points. The largest gains occur on counting questions, with improvements of 10.20 and 12.61 points, respectively. These results show that selective reasoning improves both accuracy and inference efficiency.