Geometric 4D Stitching: Moving Beyond the Optimization Bottleneck in 4D Generation
Geometric 4D Stitching for Grounded 4D Generation
The paper introduces Geometric 4D Stitching (G4S), an explicit framework for constructing expandable 4D scene representations from sparse generative videos. It replaces expensive radiance-field optimization with a geometry-aware approach that refines and stitches newly revealed regions into a consistent 4D mesh. The method achieves high-quality 4D scene expansion in under 10 minutes on a single GPU.
TL;DR
Geometric 4D Stitching (G4S) is a new framework that enables fast, consistent, and editable 4D scene generation. Unlike traditional methods that struggle with "inpainting collapse" and expensive radiance-field optimization, G4S uses a surgical approach: it identifies missing geometric regions and "stitches" in newly generated content that has been geometrically anchored to the source. It achieves SOTA results in under 10 minutes on a single RTX 5090.
The Problem: The Ill-Posed Nature of Generative 4D
Traditional 4D generation pipelines follow a standard ritual:
- Take a sparse video.
- Use a Video Diffusion Model to generate thousands of "dense" views.
- Optimize a NeRF or Gaussian Splatting model to fit these views.
The catch? This process is fundamentally ill-posed. Current Novel-View Synthesis (NVS) models are "visually greedy"—they prioritize looking good over being 3D-consistent. When you feed these inconsistent views into a radiance-based optimizer, the model gets confused. Instead of a clean 3D structure, it learns "view-dependent hacks" (like shimmering artifacts or duplicated volumes) to explain the conflicting data.
Methodology: The Surgeon’s Approach to 4D
Instead of trying to "fix" the whole scene through dense optimization, the authors of G4S propose a Region-Level Geometric Completion strategy.
1. Identifying the Gaps
The framework doesn't generate a whole new video. It calculates the Information-Addition Region. By projecting a raw mesh and comparing it to a point cloud, the system can pinpoint exactly where "silhouette curtains" (holes created by occlusion) occur. These are the only areas where the generative model is allowed to provide new data.
Figure 1: The G4S Pipeline—from initial mesh construction to geometric refinement and final stitching.
2. Pyramidal Geometric Alignment
This is the "secret sauce." Since generated depth from NVS backbones (like DA3) rarely aligns perfectly with the original source geometry, G4S introduces a Pyramidal Refinement module. It estimates a local scale-shift field that gradually aligns the new "stitch" to the existing "anchor" geometry. This ensures the new geometry doesn't "float" or "clip" through the existing scene.
Figure 2: Coarse-to-fine depth refinement ensures global alignment and local geometric detail preservation.
3. Explicit Stitching
Finally, the refined candidate pixels are back-projected and added to the mesh. Because the representation is explicit (vertices and faces) rather than implicit (radiance fields), it is inherently more stable and supports downstream tasks like object removal or scene modification.
Experimental Battleground
The results are striking. Across metrics like VBench (perceptual quality) and ATE (trajectory accuracy), G4S consistently beats established baselines like D-NeRF and 4DGS.
| Metric | D-NeRF | 4DGS | Ours (G4S) |
|---|---|---|---|
| Subject Consistency ↑ | 0.796 | 0.844 | 0.937 |
| Image Quality ↑ | 0.457 | 0.381 | 0.708 |
| ATE Mean (Motion) ↓ | 0.135 | 0.033 | 0.019 |
Figure 3: Qualitative comparison showing G4S producing far more coherent novel-view renderings than radiance-based baselines.
Conclusion: A Shift in Paradigm
Geometric 4D Stitching proves that for 4D generation, less is more. By focusing on sparse, grounded geometric updates rather than dense, unconstrained optimization, the authors have unlocked a path toward 4D content that is not only fast but functionally editable.
Limitations: The method still depends on the internal capacity of the NVS backbone. If the foundation model fails to hallucinate plausible content for a massive gap, the stitch will reflect that failure. However, as foundation models improve, G4S provides the perfect "glue" to turn those images into persistent 4D worlds.
