I am not a fan of this benchmark, nor the interpretation of Simon's. Can you draw a pelican riding a bike, and that would pass with flying colors if ranked by a diverse set of human judges? If not, you have your answer r.e. test credibility.
How are ergonomics compared to pytorch, though? Adoption can be also driven by frictionless research (e.g. torch vs. tf comes to mind). Repo is missing proper docs aimed at early adopters imho