If you look at the profiler screenshot at the bottom of the page, the time spent in the cuRAND kernel (cyan) is ~2x longer than in the numbapro kernel (purple). Even if the numbapro kernel is written in CUDA-C and suppose it will be a lot faster, you still won't hit 100x speedup. In addition, all kernels are double-precision, the 100x speedup is more common for single-precision kernels.