StyleForge:基于超图场中反事实推理的室内家具风格生成
StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
AI总结:
StyleForge是基于动态超图风格场的室内家具风格生成框架,通过反事实风格偏好学习解决固定布局家具的风格冲突,在3D-FRONT数据集上实现了优于基线的家具检索和场景风格一致性。
AI中文摘要:
固定布局的室内家具风格生成需要选择资产以构成协调的房间,且不得改变规定的家具类别、位置、朝向或尺寸。现有方法通常独立检索每个资产或依赖静态局部关系,在场景组合后易出现形状、材质和颜色冲突。我们引入StyleForge,这是一个基于动态超图风格场的场景级结构化选择框架。冻结的多模态大语言模型从开放式风格请求和固定布局中提取结构化风格先验,而StyleForge为每个家具槽维护可学习的候选分布。在目标风格的条件下,动态超图风格场自适应激活并加权布局诱导的超边,以捕捉家具间的高阶依赖关系。随后的反事实风格偏好学习将每个候选视为当前风格场中的局部替换,并用马氏能量评估其上下文兼容性。训练在优化风格场和候选对数之间交替进行。推理时,模型保持冻结,测试时训练仅更新特定房间的候选对数,随着全局场景上下文的演变逐步纠正跨槽风格冲突。在3D-FRONT上的实验表明,该方法实现了最先进的家具检索和场景级风格一致性,比对象级和场景级检索基线产生更协调的固定布局家具布置。
英文摘要:
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.