arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HounsWorld:用于隐藏患者状态读取、重建与模拟的多模态世界模型

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation

Yunhao Bai, Zhongwei Qiu, Guangyu Guo, Yiming Huang, Tony C. W. Mok, Qinji Yu, Ling Zhang, Yan Wang

arXiv 2608.12904首次发表:更新:

发表机构

East China Normal University; DAMO Academy, Alibaba Group; Hupan Laboratory; Zhejiang University(华东师范大学; 阿里巴巴集团达摩院; 湖畔实验室; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出3B参数多模态世界模型HounsWorld,结合CT扫描与临床语言,通过共享潜在患者状态实现读取、重建、模拟三类任务,在HounsBench基准上表现优异,提升了CT理解能力。

AI 中文摘要

临床智能需从不完整观测中估算患者的潜在状态,而非学习从扫描到答案的孤立映射。容积医学图像提供解剖结构、衰减系数及病灶的密集观测,而临床语言则提供稀疏但互补的语义观测。我们将以CT为核心的智能定义为对共享潜在患者状态的推理,在此框架下,读取、重建与模拟均成为依赖状态的预测问题。为实现该观点,我们引入了HounsBench——一个以计算机断层扫描(CT)为核心的患者状态基准,其通过患者不相交划分及各任务家族的指标,统一了这三类任务;同时提出HounsWorld,这是一个3B参数的多模态世界模型,通过联合理解-生成学习将容积扫描与语言视为共享状态的观测。共享Transformer形成隐式患者状态估计,支持三类输出:读取状态的查询条件答案、以语言重建状态的报告与描述,以及用于低剂量去噪、虚拟对比增强、解剖结构约束的文本-掩码到容积生成的特定条件CT容积。零初始化CT适配器保留预训练多模态映射,而条件明确的亨斯菲尔德单位(Hounsfield-unit)窗口采样则暴露具有临床意义的密度观测。HounsWorld在所有三类任务中均表现出优异性能,同时通过临床结构化补全持续提升CT理解能力。本项目可在该https URL获取。

英文摘要

Clinical intelligence requires estimating a patient's underlying condition from incomplete observations rather than learning isolated mappings from scans to answers. Volumetric medical images provide dense observations of anatomy, attenuation, and lesions, whereas clinical language provides sparse but complementary semantic observations. We formulate CT-centered intelligence as inference over a shared latent patient state, under which readout, reconstruction, and simulation all become state-dependent prediction problems. To operationalize this view, we introduce HounsBench, a computed tomography (CT) centric patient-state benchmark that unifies these three task families with patient-disjoint splits and per-family metrics, and HounsWorld, a 3B multimodal world model that treats volumetric scans and language as observations of the shared state through Joint Understanding-Generation Learning. A shared transformer forms an implicit patient-state estimate and supports three outputs: query-conditioned answers that read out the state, reports and captions that reconstruct it in language, and condition-specific CT volumes for low-dose denoising, virtual contrast enhancement, and anatomy-constrained text-and-mask-to-volume generation. Zero-initialized CT adapters preserve pretrained multimodal mappings, while condition-explicit Hounsfield-unit window sampling exposes clinically meaningful density observations. HounsWorld shows strong performance across all three task families while consistently improving CT understanding through clinically structured completion. Our project is available at https://github.com/byhwhite/HounsWorld.git

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑