发表机构
University of Pittsburgh(匹兹堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对临床步态分析中黑盒模型不可解释及VLM直接应用易产生幻觉的问题,提出无需训练的DrGait主体框架,通过分诊-验证-综合工作流解耦语义推理与几何感知,借助确定性生物力学工具验证假设,减少幻觉并生成透明可审计的临床报告。
AI 中文摘要
当前的临床自动化步态分析依赖于不可解释的黑盒分类器。尽管视觉-语言模型(VLM)提供了强大的推理能力,但直接将其应用于步态视频往往会导致幻觉,因为它们难以从原始视觉上下文中测量细微的几何偏差。为解决这一问题,我们提出了DrGait,一个无需训练的主体框架,将VLM的角色从直接视觉推理者转变为临床规划者。DrGait通过结构化的分诊-验证-综合(TVS)工作流将语义推理与几何感知解耦。给定输入视频和一组基本时空指标,DrGait主体首先执行启发式分诊以提出诊断假设,然后通过自主调用确定性生物力学工具进行验证,这些工具基于重建的3D网格轨迹、分割的2D姿态轨迹和以事件为中心的视频证据运行。最后,闭环机制根据反馈递归更新主体的推理上下文。通过将VLM的推理锚定在可验证的几何和时间测量上,DrGait减少了幻觉,在生成透明且可审计的临床报告的同时实现了具有竞争力的诊断准确性。
英文摘要
Current automated gait analysis for clinical applications relies on uninterpretable black-box classifiers. Although Vision-Language Models (VLMs) offer strong reasoning capabilities, applying them directly to gait videos often leads to hallucinations, because they struggle to measure subtle geometric deviations from raw visual contexts. To address this, we introduce DrGait, a training-free agentic framework that shifts the VLM's role from a direct visual reasoner to a clinical planner. DrGait decouples semantic reasoning from geometric perception through a structured Triage-Verification-Synthesis (TVS) workflow. Given an input video and a set of basic spatiotemporal metrics, the DrGait agent first performs a heuristic triage to propose diagnostic hypotheses, which are then verified by autonomously calling deterministic biomechanical tools that operate on reconstructed 3D mesh trajectories, segmented 2D pose tracks, and event-centered video evidence. Finally, a closed-loop mechanism recursively updates the agent's reasoning context based on the feedback. By anchoring VLM's reasoning in verifiable geometric and temporal measurements, DrGait reduces hallucinations, achieving competitive diagnostic accuracy while generating transparent and audit-ready clinical reports.
Comments76 pages, 6 figures