Are you saying that cloud models are not verifying the draft model predictions? The way draft models are used in something like llama.cpp results in exactly zero degredation of output quality, with the larger model verifying each draft model token and discarding it if it does not match.