大型语言模型 (LLMs) 在曼哈顿环境的 3D 室内合成中表现出了卓越的能力。然而,, 现有方法通常无法捕获非曼哈顿环境中的合理对象布局模式,,主要是因为它们很难对非正交空间关系进行建模,,从而导致较高的几何违规和较低的物理保真度。为了应对这一挑战,,我们提出了 SPG-Layout, 一种新颖的文本驱动框架,旨在在复杂的非曼哈顿环境中生成物理上合理的室内场景。具体来说,我们首先利用对象分布的统计先验来指导训练过程,增强环境理解和保真度。此外, 反映了人类设计工作流程, 我们采用分层布局策略,优先考虑大型对象的放置, 从而大大减少布局违规。通过协同这些组件, SPG-Layout 实现了语义真实性和物理合理性的平衡优化。为了评估这些复杂环境中的性能,,我们构建了一个包含 500 个不同的非曼哈顿环境的新基准。大量实验表明,SPG-Layout 在曼哈顿和非曼哈顿环境中始终显着优于现有方法。该代码将公开发布。

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge, we propose SPG-Layout, a novel text-driven framework designed to generate physically plausible indoor scenes within complex non-Manhattan environments. Specifically, we first utilize statistical priors of object distributions to guide the training process, enhancing environmental understanding and fidelity. Furthermore, mirroring human design workflows, we adopt a hierarchical layout strategy that prioritizes the placement of large objects, thereby substantially minimizing layout violations. By synergizing these components, SPG-Layout achieves a balanced optimization of semantic realism and physical plausibility. To evaluate performance in these complex settings, we constructed a new benchmark comprising 500 diverse non-Manhattan environments. Extensive experiments demonstrate that SPG-Layout consistently and significantly outperforms existing methods across both Manhattan and non-Manhattan environments. The code will be publicly released.

科目: 人工智能 (cs.AI); 计算机视觉和模式识别 (cs.CV)

Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)