Wait until you see the same concept combined with NeRF idea. The output won’t be 3d shapes but another model that can generate realistic and geometrically consistent images of a scene viewed from different angles.
Patent what? Supervised sound localization is not novel. There exist myriads of published work on that topic. The only novelty I see in this work is the similarity between performance of their trained model and that in humans.