Do we actually need GPUs to run this? There is no training involved, only inference, and CPUs (or low-end GPUs) should comfortably run the workload, at least for a couple of faces.
any lossy video-compression algorithm face the same challenge. what you are seeing is artificial and is constructed by the algorithm to minimize the perceptual difference between the real and the constructed video feed.