you could try but the amount of angle information in the data will be so small that noise will destroy the results. Said another way, the diffuser makes the forward model more invertible (better condition number).
These are not just surface voxels, but occlusions will limit angle info. In the paper, it's explained how it is possible to extract 100 Million voxels of information out of 1 Million pixels measured - by using compressed sensing and assuming that the object has some structure (is not totally random 3D distribution). This compressed sensing trick is not new but is a large part of what distinguishes this work from light field cameras, which cannot extract so much information and so must trade resolution for depth info.
Actually, lenslets/Hartmann mask are not shift-invariant, and getting depth info (3D) requires loss of shift invariance (see http://web.mit.edu/2.717/www/stein.pdf). Lenslet-based light field cameras also have very extensive calibration routines to align the lenslets on a per-pixel basis.