Faithfully capturing the intricate relationship between motion and geometry in dynamic scenes is essential for 4D reconstruction. Recent state-of-the-art methods rely on monolithic deformation networks or global time-basis factorizations that struggle to represent complex, non-rigid topological changes. We propose MoBa-GS, a framework that resolves this entanglement by introducing a structural inversion: a spatially-factorized motion field coupled with adaptive geometric optimization. First, our model learns a low-frequency dynamic canonical space to represent coarse scene motion. Next, it decomposes complex, non-rigid motion into a Spatially-Varying Motion Basis of local kinematics, predicted from the canonical geometry, which is then linearly combined using dynamic blending weights. This formulation directly acts as an implicit neural scaffold, recovering both geometry and motion from random initialization, thereby removing the dependency on SfM point cloud priors. This design is further augmented with motion-guided densification and positional annealing to reduce geometry overfitting. Extensive experiments show that our framework surpasses prior state-of-the-art methods in reconstruction fidelity. Enabled by time-invariant caching, MoBa-GS requires an order-of-magnitude shorter training time (~18 minutes), a compact storage (~11 MB), and achieves real-time rendering speeds (~163 FPS). Our work establishes a new foundation for high-fidelity, efficient 4D representations without relying on explicit geometric priors.
Figure 1: MoBa-GS Overview. (Stage 1) A low-frequency dynamic canonical space is learned to capture coarse, global scene motion. (Stage 2) A Spatially-Varying Motion Basis decomposes the residual non-rigid motion into local kinematic subspaces, one per Gaussian, predicted from the canonical geometry. Dynamic blending weights linearly combine the basis vectors into a full per-Gaussian deformation field. (Training) Motion-guided densification and positional annealing progressively refine geometry, removing the need for SfM initialization. At test time, time-invariant feature caching decouples rendering from network evaluation, achieving >163 FPS.
Novel view synthesis comparisons on D-NeRF and NeRF-DS datasets.
Table 1 — D-NeRF Synthetic Dataset
Comparison on D-NeRF synthetic dataset (PSNR / SSIM / LPIPS). MoBa-GS surpasses prior art in reconstruction fidelity while delivering an order-of-magnitude reduction in both training time and storage.
Table 2 — NeRF-DS Real-World Dataset
Mean comparison across all NeRF-DS scenes. MoBa-GS handles real-world reflective and non-rigid surfaces without SfM priors.
Figure 2: D-NeRF qualitative results. MoBa-GS recovers sharper boundaries and reduces floaters in fast-moving regions compared to SC-GS, Deformable-GS, and 4D-GS.
Figure 3: NeRF-DS qualitative results. The Spatially-Varying Motion Basis captures per-region motion heterogeneity on real-world reflective and deformable surfaces with high fidelity.
Figure 4: Ablation studies. We ablate the dynamic canonical space, the Spatially-Varying Motion Basis, motion-guided densification, and positional annealing. Each component provides measurable, complementary improvements in both fidelity and efficiency.
@inproceedings{tan2026mobags,
author = {Tan, Guan Yuan and Pal, Arghya and Rajanala, Sailaja
and Phan, Rapha{\"{e}}l C-W and Ting, Chee-Ming},
title = {MoBa-GS: Learning a Spatially-Varying Motion Basis
over a Dynamic Canonical Space for 4D Reconstruction},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026},
}