Structured Neural Representations for 3D Shape and Scene Editing
Description
The generation of 3D assets from scratch has become a major area of interest in computer vision and graphics research, driven by applications in several visual mediums, including games, animation, and virtual reality. Despite rapid progress in generative modeling, modifying existing assets in a controllable and interpretable manner remains difficult. Precise 3D manipulation is fundamentally a representation problem: the types of edits a neural system can perform depend on how the scene is processed internally. This dissertation studies how different forms of structure influence the quality, controllability, interpretability, and expressiveness of supported editing operations.
We begin in the second chapter by exploring early representations for generative 3D models, focusing on using tetrahedra as an alternative to dense voxel-based shape synthesis. By training standard variational autoencoders over datasets of 3D assets represented using this discretization, our method learns coherent latent shape embeddings that support shape editing via latent interpolation and arithmetic. Both surface and volumetric geometry can be extracted directly from the learned tetrahedral grid, producing explicit mesh assets suitable for downstream tasks.
However, this class of method is fundamentally limited. Learning meaningful latent embeddings requires training over large collections of 3D assets, while large-scale 3D datasets remain substantially less abundant than image data. Furthermore, edits performed through latent manipulation provide only coarse global control over generated geometry, making precise semantic modifications difficult to specify. These limitations motivated the development of editing approaches that leverage rich semantic information encoded in large multimodal models pretrained on readily available image-text data.
In the third chapter, we study text-based methods for optimizing 3D geometry using natural language as an interface for specifying edits. A common approach in this vein is to directly optimize mesh vertices and differentiably render the resulting shape to obtain image-text similarity gradients, but this often leads to noisy and unstable deformations. To address this limitation, we adopt an alternative parameterization of the deformation field on vertices that produces smoother and more semantically coherent changes.
Despite our method producing smoother and more coherent deformations, optimization-based editing methods are difficult to control in practice. Direct deformation of explicit geometry also restricts the range of allowable modifications. In the fourth chapter, we address these limitations using a feed-forward masked editing architecture that conditions generation on both the initial input geometry and localized edit specifications. Our method edits in a single forward pass rather than through iterative optimization, producing faster results. The edits are also more predictable as the model directly incorporates the specified edit into the generated 3D geometry, while preserving unedited regions of the original shape.
In the fifth chapter, we shift from editing individual shapes and study structured scene representations for editable neural rendering. This work focuses on recovering semantically meaningful objects and parts within scenes by lifting image-space information into 3D representations. We investigate methods for converting multi-view semantic information into coherent 3D scene decompositions, using either supervised annotations or unsupervised feature affinity methods. The resulting representations enable individual scene components to be isolated and manipulated through operations such as translation, scaling, and scene recomposition.
Files
PhD_Dissertation.pdf
Files
(39.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:1ef05c2831fa2f94000432f6cc03ee51
|
39.6 MB | Preview Download |