ECCV 2026

MoBa-GS: Learning a Spatially-Varying Motion Basis
over a Dynamic Canonical Space for 4D Reconstruction

Guan Yuan Tan, Arghya Pal, Sailaja Rajanala, Raphaël C.-W. Phan, Chee-Ming Ting
School of Information Technology, Monash University Malaysia
24.14 dB
Mean PSNR on NeRF-DS
(state of the art)
13 min
Training time
(vs ~3 h for D-MiSo)
163 FPS
Real-time rendering
(time-invariant caching)
11 MB
Storage footprint
(~39k Gaussians)
SfM-Free
Random initialization
no COLMAP required

MoBa-GS resolves the entanglement of motion and geometry with a structural inversion: instead of factorizing motion over time with a globally shared trajectory dictionary, we factorize over space — every Gaussian predicts its own local motion basis from its canonical position. Because that basis is time-invariant it can be cached, giving real-time rendering at 163 FPS in ~11 MB after 13 min of training — all from random initialization, with no SfM point cloud.

Presentation

Five-minute overview of the method and results.

Video unavailable in your region? Download the MP4 (8 MB).

Abstract

Faithfully capturing the intricate relationship between motion and geometry in dynamic scenes is essential for 4D reconstruction. Recent state-of-the-art methods rely on monolithic deformation networks or global time-basis factorizations that struggle to represent complex, non-rigid deformation. We propose MoBa-GS, a framework that resolves this entanglement by introducing a structural inversion: a spatially-factorized motion field coupled with adaptive geometric optimization. First, our model learns a low-frequency dynamic canonical space to represent coarse scene motion. Next, it decomposes complex, non-rigid motion into a Spatially-Varying Motion Basis of local kinematics, predicted from the canonical geometry, which is then linearly combined using dynamic blending weights. This formulation acts as an implicit neural scaffold, recovering both geometry and motion from random initialization and thereby removing the dependency on SfM point cloud priors. The design is further augmented with motion-guided densification and positional annealing to reduce geometry overfitting. Extensive experiments show that our framework surpasses prior state-of-the-art methods in reconstruction fidelity while training in roughly thirteen minutes and rendering in real time from an eleven-megabyte model.

Method

MoBa-GS architecture

Figure 1: Overview. A Base Motion Network learns a low-frequency prior that deforms canonical positions into a dynamic canonical space. The residual motion is then factorized: a Basis Predictor maps the static canonical position to a local motion basis V ∈ ℝ3×K, while a Weight Predictor maps the dynamic state to blending weights. Because V never depends on t, it is pre-computed and cached once the geometry stabilizes, reducing deformation at render time to a dot product.

Learned motion basis and predicted weights

Figure 2: The learned basis is interpretable. A static, time-invariant toolkit of local motion primitives (middle) is combined with weights that change over time (bottom) to match the ground-truth motion (top). Because the basis is a function of geometry alone, it doubles as an unsupervised map of which regions are rigid and which deform.

Qualitative Comparisons

Novel-view synthesis on monocular real-world video. Each clip is a 2×2 grid — ground truth and MoBa-GS on top, baselines below.

NeRF-DS — Basin Reflective · thin structure
NeRF-DS — Plate Fast motion
HyperNeRF — Chicken Extreme non-rigid deformation

Played at reduced speed. The deforming region is where the gap is clearest — baselines lose the object boundary during fast motion.

Quantitative Results

NeRF-DS — mean over 7 scenes
MethodPSNR ↑SSIM ↑LPIPS ↓
SC-GS22.250.8240.203
4DGS22.540.8370.212
Motion-GS23.710.8310.240
Deformable-GS23.760.8480.180
D-MiSo23.900.8510.151
MoBa-GS (Ours)24.140.8570.197

D-MiSo attains a lower LPIPS; we lead on PSNR and SSIM.

HyperNeRF — mean over 4 scenes
MethodPSNR ↑SSIM ↑
SC-GS20.920.63
MotionGS20.960.50
Deformable-GS22.060.59
D-MiSo22.470.62
MoBa-GS (Ours)22.730.63

LPIPS omitted following the protocol of MotionGS.

Efficiency — NeRF-DS, single NVIDIA A100
MethodTraining time
D-MiSo~3 hours
MoBa-GS (Ours)13 min

MoBa-GS renders at 163 FPS from ~39k Gaussians in 11.3 MB. Both methods timed on the same single NVIDIA A100.

Learned Motion Basis

Learned spatially-varying motion basis

Figure 3: Spatially-varying basis vectors. Basis vectors for the most rigid point (left) and the most non-rigid point (right) in the canonical space of the ‘As’ scene. The rigid point’s basis is well distributed, indicating capacity for general transformations; the non-rigid point’s is directionally aligned, a specialized representation for localized deformation.

BibTeX

@inproceedings{tan2026mobags,
  title     = {MoBa-GS: Learning a Spatially-Varying Motion Basis over a
               Dynamic Canonical Space for 4D Reconstruction},
  author    = {Tan, Guan Yuan and Pal, Arghya and Rajanala, Sailaja and
               Phan, Rapha{\"e}l C.-W. and Ting, Chee-Ming},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}