arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VeriScene:通过世界模型智能体从法律证据重建犯罪现场

VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent

Kevin Chuanpu Fu, Yongsen Zheng, Zee Kin Yeong, Kwok-Yan Lam

arXiv 2609.08342首次发表:更新:

发表机构

Nanyang Technological University; Singapore Academy of Law(南洋理工大学; 新加坡法律学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VeriScene利用世界模型智能体,通过审计循环融合法医照片和证人陈述,重建犯罪现场并生成重现视频,在基准上显著提升事实一致性和时间连贯性。

AI 中文摘要

世界模型以文本、照片和图表等多模态输入,依据物理定律生成动态场景,从而开启了一个引人注目的应用:融合多模态法律证据以重建犯罪现场,并重演犯罪可能如何实施。然而,将原始、无组织的证据直接输入世界模型在法医应用中会失败:它会悄然丢弃证据,对相互矛盾的证词含糊其辞,并产生违反证据记录的运动。本文提出了VeriScene,一个编排世界模型的智能体:它从法医照片和可靠性不同的证人陈述中重建犯罪现场,使每项主张都可追溯到证据,每个运动在物理上合理。VeriScene在审计循环下迭代地将证据融合成带引用的叙述,通过在世界模型中进行探针回放并注入纠正性约束来验证假设的动态,并从融合的关键帧将犯罪渲染为重现视频。在一个包含7种物理驱动案件类型(139张法医风格照片和65份带有植入不可靠性的陈述)的25个犯罪场景基准上,VeriScene在20个测试场景中达到了0.9014的证据覆盖率和0.7217的事实一致性(0-1量表),在事实一致性上比端到端多模态LLM基线高出20.35%,在时间连贯性上高出34.88%,同时在四个LLM编排后端上以每场景1.82美元的成本实现了泛化。

英文摘要

World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance with the laws of physics, thus opening a compelling application: fusing multimodal legal evidence to re-create a crime scene and re-enact how an offence could have been committed. However, feeding the raw, unorganized evidence into a world model fails in forensic use: it silently drops evidence, glosses over contradictory testimony, and produces motion that violates the evidentiary record. This paper presents VeriScene, an agent that orchestrates the world model: it reconstructs crime scenes from forensic photographs and witness statements of varying reliability, keeping every claim traceable to evidence and every motion physically plausible. VeriScene iteratively fuses the evidence into a cited narrative under an auditing loop, verifies the hypothesized dynamics via probe rollouts in the world model with corrective constraint injection, and renders the offence as a re-enactment video from a fused keyframe. On a benchmark of 25 crime scenarios across 7 physically-driven case types (139 forensic-style photographs and 65 statements with planted unreliability), VeriScene attains 0.9014 evidence coverage and 0.7217 factual consistency (0-1 scale) on the 20 test scenes, outperforming an end-to-end multimodal-LLM baseline by 20.35% in factual consistency and 34.88% in temporal coherence, while generalizing across four LLM orchestration backends at USD 1.82 per scene.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑