An MLP-based NeRF actually has a comparable number of parameters to plenoxels (it's not 2-3 orders of magnitude smaller). The original NeRF is 8 dense layers with 256 channels and then this one adds another network with 6 dense layers with 64 channels, so roughly speaking 7x256x256 + 5x64x64. And remember that the voxel grid is sparse, though we don't get exact numbers here. We shouldn't be miserly with our megabytes in 2021. What concerns me is how HyperNeRF requires 64 hours of training time on 4 TPU v4s; if you want to use this for communication or entertainment, it's light-years away from interactive.
Extending plenoxels to support dynamic objects would be great future work.
Extending plenoxels to support dynamic objects would be great future work.