Cloud TPU Pods Break AI Training Records(cloud.google.com)
cloud.google.com
Cloud TPU Pods Break AI Training Records
https://cloud.google.com/blog/products/ai-machine-learning/cloud-tpu-pods-break-ai-training-records
4 comments
Disclosure: I work on Google Cloud.
I know. Within MLperf, to reach consensus, it was agreed that “dollars are great, but also game-able via price cuts, lets normalize to a reference system”. It’s not the largest box though, it’s one that everyone was able to provide, so that it lets you compare through one level of indirection.
I don’t agree with it, because it makes it so hard to compare to funky hardware (“Should I do inference on this chip from < startup >? I dunno, what’s it going to cost?”). Having to keep going back to some reference box is undesirable, but there’s nothing stopping someone from making a spreadsheet equivalent that translates these to dollars on the cloud providers (harder for on-prem submissions which reopens the “rent vs buy” debate).
I know. Within MLperf, to reach consensus, it was agreed that “dollars are great, but also game-able via price cuts, lets normalize to a reference system”. It’s not the largest box though, it’s one that everyone was able to provide, so that it lets you compare through one level of indirection.
I don’t agree with it, because it makes it so hard to compare to funky hardware (“Should I do inference on this chip from < startup >? I dunno, what’s it going to cost?”). Having to keep going back to some reference box is undesirable, but there’s nothing stopping someone from making a spreadsheet equivalent that translates these to dollars on the cloud providers (harder for on-prem submissions which reopens the “rent vs buy” debate).
I'm not (currently) objecting to the actual MLPerf results.
I'm objecting to the fact that your marketing team (at least, I hope it was your marketing team) chose to represent 5 different GPU systems as the same color and 3 different TPU systems as another color in one bar chart.
There's a book with a chapter or two on how to reasonably compare the performance of computer systems. Given that the author (Patterson) works on TPUs, I would expect better.
I'm objecting to the fact that your marketing team (at least, I hope it was your marketing team) chose to represent 5 different GPU systems as the same color and 3 different TPU systems as another color in one bar chart.
There's a book with a chapter or two on how to reasonably compare the performance of computer systems. Given that the author (Patterson) works on TPUs, I would expect better.
Thanks for the feedback. As more and more new ML system architectures emerge over the next few years, these performance comparisons will get even more complicated. Chip-to-chip comparisons and server-to-server comparisons are arbitrary, and it will sometimes be difficult to define these boundaries at all.
Our impression is that top-line performance comparisons independent of system size (like the comparison in the blog post) and performance-per-dollar comparisons are the easiest to understand, but we're certainly open to other ideas.
Our impression is that top-line performance comparisons independent of system size (like the comparison in the blog post) and performance-per-dollar comparisons are the easiest to understand, but we're certainly open to other ideas.
> top-line performance comparisons independent of system size
Then do that, just don't do it all in one chart with the systems only specified in the fine print.
Then do that, just don't do it all in one chart with the systems only specified in the fine print.
Author of the blog post here.
Cloud TPUs are designed to maximize performance-per-dollar, so you are right that pure performance comparisons at maximum scale don't tell the whole story.
The most straightforward performance-per-dollar comparisons would be among several different hardware configurations across the major public clouds. However, we haven't yet seen any other public cloud MLPerf submissions at scales comparable to Cloud TPU Pods, so there isn't currently a strong baseline available for comparison. It's also not clear whether public cloud networking will ultimately be able to match the performance of the network hardware that was used to produce the largest-scale on-premise MLPerf submissions.
Cloud TPUs are designed to maximize performance-per-dollar, so you are right that pure performance comparisons at maximum scale don't tell the whole story.
The most straightforward performance-per-dollar comparisons would be among several different hardware configurations across the major public clouds. However, we haven't yet seen any other public cloud MLPerf submissions at scales comparable to Cloud TPU Pods, so there isn't currently a strong baseline available for comparison. It's also not clear whether public cloud networking will ultimately be able to match the performance of the network hardware that was used to produce the largest-scale on-premise MLPerf submissions.
Can you elaborate on the performance-per-dollar part? IE, I'm seeing GCP providing V100 at $2.48/hour and a TPUv3 at $8.00/hour
Sure. The $2.48/hour per V100 GPU on GCP does not include the price of the CPU host; that is purely the price to rent a single accelerator. By contrast, a network-attached Cloud TPU v3 device includes both a CPU host and four connected TPU v3 chips that collectively deliver up to 420 teraflops. Furthermore, each individual V100 GPU on GCP has 16 GB of memory, whereas the Cloud TPU v3 device has 128 GB of HBM.
The best apples-to-apples performance-per-dollar comparison we have publicly available was published last fall, and it compared the performance and cost of using various Cloud TPU v2 Pod slice sizes with the performance and cost of using various numbers of V100 GPUs attached to a single GCP host:
https://cloud.google.com/blog/products/ai-machine-learning/n...
We went to great lengths to ensure that we trained exactly the same version of ResNet-50 to the same accuracy in the same way across all hardware configurations. The methodology predated MLPerf and is documented in full here:
https://github.com/tensorflow/tpu/blob/master/benchmarks/Res...
If you were going to do a similar performance-per-dollar comparison today, the simplest approach might be to try to get the code from NVIDIA's MLPerf 0.6 submissions running at scale on one or more major public clouds using the fastest-available networking technology that each cloud provides:
https://github.com/mlperf/training_results_v0.6/tree/master/...
It would be very interesting to see how distributed training performance using large-scale GPU clusters in public clouds compares with the published on-premise MLPerf performance numbers using exactly the same MLPerf code and methodology. With these measurements in hand, it would then be straightforward to make performance-per-dollar comparisons with Cloud TPU v3 Pod slices of various sizes.
The best apples-to-apples performance-per-dollar comparison we have publicly available was published last fall, and it compared the performance and cost of using various Cloud TPU v2 Pod slice sizes with the performance and cost of using various numbers of V100 GPUs attached to a single GCP host:
https://cloud.google.com/blog/products/ai-machine-learning/n...
We went to great lengths to ensure that we trained exactly the same version of ResNet-50 to the same accuracy in the same way across all hardware configurations. The methodology predated MLPerf and is documented in full here:
https://github.com/tensorflow/tpu/blob/master/benchmarks/Res...
If you were going to do a similar performance-per-dollar comparison today, the simplest approach might be to try to get the code from NVIDIA's MLPerf 0.6 submissions running at scale on one or more major public clouds using the fastest-available networking technology that each cloud provides:
https://github.com/mlperf/training_results_v0.6/tree/master/...
It would be very interesting to see how distributed training performance using large-scale GPU clusters in public clouds compares with the published on-premise MLPerf performance numbers using exactly the same MLPerf code and methodology. With these measurements in hand, it would then be straightforward to make performance-per-dollar comparisons with Cloud TPU v3 Pod slices of various sizes.
Probably Google is still doing that 1 "Cloud TPUv3" = 4 TPUv3 chips.
I cannot stand Google people's tendency to explain or sometimes rebuttal comments.
You are talking to your customers who is paying or considering paying, or in search of products.
Take the feedback, if it can be done, and it's beneficial, do it and report so.
Or stop explaining... That's simply not professional for a cloud provider...
You are talking to your customers who is paying or considering paying, or in search of products.
Take the feedback, if it can be done, and it's beneficial, do it and report so.
Or stop explaining... That's simply not professional for a cloud provider...
Why performance per dollar over performance per watt?
Performance per watt doesn't care where the systems are (on-prem vs cloud) and should track perf/$ (unless perf/$ is mostly determined by subsidies).
Performance per watt doesn't care where the systems are (on-prem vs cloud) and should track perf/$ (unless perf/$ is mostly determined by subsidies).
I think hardware acquisition costs can dominate GPU prices, so perf/watt might not be a realistic measure for the cost of running such large scale/high performance experiments.
For instance, the DGX-2 has an MSRP of $399,000 and consumes 10 kW of power [1]. The average commercial electricity cost across the US is about $0.11/kWh [2], so a DGX-2 running at full tilt costs $1.10 an hour in electricity. Thus, a DGX-2 running at 100% utilization for 3 years costs $9,636 in electricity, which is ~2.5% of the cost of the box itself.
Of course, you probably could get DGX-2s for a lot cheaper if you are buying 100 of them, but the acquisition costs are still going to be significant vis-a-vis power costs.
[1]: https://www.anandtech.com/show/12587/nvidias-dgx2-sixteen-v1...
[2]: https://www.pacificpower.net/about/rr/cpc.html
For instance, the DGX-2 has an MSRP of $399,000 and consumes 10 kW of power [1]. The average commercial electricity cost across the US is about $0.11/kWh [2], so a DGX-2 running at full tilt costs $1.10 an hour in electricity. Thus, a DGX-2 running at 100% utilization for 3 years costs $9,636 in electricity, which is ~2.5% of the cost of the box itself.
Of course, you probably could get DGX-2s for a lot cheaper if you are buying 100 of them, but the acquisition costs are still going to be significant vis-a-vis power costs.
[1]: https://www.anandtech.com/show/12587/nvidias-dgx2-sixteen-v1...
[2]: https://www.pacificpower.net/about/rr/cpc.html
When comparing different hardware configurations within or between public clouds, power measurements generally aren't available, whereas prices or realistic price estimates generally are.
You might want to check the DAWN benchmark, it does just that:
[1] https://dawn.cs.stanford.edu/benchmark/index.html#imagenet-t...
[1] https://dawn.cs.stanford.edu/benchmark/index.html#imagenet-t...
Thanks, I've seen dawnbench, and it's definitely better than this presentation.
Unfortunately, their cost numbers aren't actual cost, as they don't allow using AWS spot instance pricing (which is what fast.ai does).
https://www.fast.ai/2018/04/30/dawnbench-fastai/
Unfortunately, their cost numbers aren't actual cost, as they don't allow using AWS spot instance pricing (which is what fast.ai does).
https://www.fast.ai/2018/04/30/dawnbench-fastai/
It was decided to disallow spot pricing and preemptible VMs because it’s super variable. On the day of a CVPR submission or something, you’re not going to pay the lowest rate (spot) nor perhaps be likely to get the VMs anyway (both spot and preemptible).
Said another way, spot pricing results aren’t as reproducible in the research sense. It’s also a similar constant factor between the main providers. We felt that once you knew relative price/perf, you can choose to do whatever economic analysis you’d prefer (e.g., maybe if you don’t have any datacenter space yourself, there is no price you’d pay for hardware on-premises, or maybe you are willing to use spot or preemptible).
Said another way, spot pricing results aren’t as reproducible in the research sense. It’s also a similar constant factor between the main providers. We felt that once you knew relative price/perf, you can choose to do whatever economic analysis you’d prefer (e.g., maybe if you don’t have any datacenter space yourself, there is no price you’d pay for hardware on-premises, or maybe you are willing to use spot or preemptible).
Divide by 3 to get pre-emptible price
Results are disappointing. 2x speed up over general purpose GPUs doesn’t justify a whole new hardware architecture.
When Moores law is finally dead and buried (e.g. 5nm), new architectures will be all that's left. Seems like a great time to start down the new architecture path.
[deleted]
Have semiconductor companies not been on the new architecture path already? It's not like they've just been doing die shrinks this whole time.
Yes, but the emphasis now is much greater than before. The old rules of turn the scaling crank that defined the industry for decades are no longer helping.
For actual practitioners, TPUs are incredible. The cost/performance combo is unmatched.
Now the real problem is, can you actually get a TPU pod in practice?
Now the real problem is, can you actually get a TPU pod in practice?
My personal experience with crnns and lstms was that cutting training latency from 48 hours to 24 hours didn’t make a difference in our progress. Some architectures took 4x longer to run.
48 to 1 hour would have allowed us to run a few experiment per work day. That would be huge.
If you told me that the next generation of TPUs could run the current 48 hour GPU training set in 1 hour, that would be something.
But 2x faster won’t make a difference. Generally, startups for a new chip need to demonstrate 100x improvement over the general purpose approach to get funding. So this seems like another Google vanity project.
48 to 1 hour would have allowed us to run a few experiment per work day. That would be huge.
If you told me that the next generation of TPUs could run the current 48 hour GPU training set in 1 hour, that would be something.
But 2x faster won’t make a difference. Generally, startups for a new chip need to demonstrate 100x improvement over the general purpose approach to get funding. So this seems like another Google vanity project.
I don't understand. Time-to-train is always a product of model and the amount of hardware you use. TPUs are fast and cheap allowing you to use more hardware for the same cost. The TPU design also makes it scale quite well.
If you really wanted to get from 48 hours to multiple experiments per work day, you could (probably) do that right now by scaling horizontally. However, the cost becomes a major issue when you do that on GPUs. This is where the TPU shines since you can scale out without cost crippling you.
I work on training on GPUs and the TPU is definitely an incredibly useful piece of hardware, not a vanity project.
If you really wanted to get from 48 hours to multiple experiments per work day, you could (probably) do that right now by scaling horizontally. However, the cost becomes a major issue when you do that on GPUs. This is where the TPU shines since you can scale out without cost crippling you.
I work on training on GPUs and the TPU is definitely an incredibly useful piece of hardware, not a vanity project.
According to the post, it’s only 2x cheaper. So for the same dollars spent, I would cut my time from 48 hours to 24.
But, TPUs are not standard, can’t be used for any other usecase. So those savings might not actually be realizable when everything is taken into account.
But, TPUs are not standard, can’t be used for any other usecase. So those savings might not actually be realizable when everything is taken into account.
That is not what the blog is saying. The blog post does not consider cost - it simply shows that a TPU v3 Pod is twice as fast as the largest DGX-2h cluster.
Why would you use cloud if you do a lot of training? Quad 2080Ti systems cost ~$7k + electricity. Assuming two such systems are equivalent in speed to 1 TPUv3 (4 chips), owning them would be more cost efficient after 4-5 months of training (depending on your electricity costs).
TPUs do have the memory capacity advantage though (over 2080Ti).
TPUs do have the memory capacity advantage though (over 2080Ti).
Is it me or are those results somewhat underwhelming if anything? Dedicated hardware for a 2x speedup at best, tossup for most results, and only competes in some categories. Not to be just a NVIDIA fan here, surely there is value in dedicated training hardware, but just surprising that benefit isn't bigger!
Disclosure: I work on Google Cloud (even with Zak sometimes).
The DGX-2h is a beast! Don’t parse this as “huh, TPU Pods are about the same as just a few V100s”. The data sheet [1] is probably the easiest to follow, but their writeup is more informative [2].
These are souped up V100s, with awesome networking, which is pretty similar in style to a TPU Pod. So I’d say that they’re both purpose built systems for distributed ML training. The name for the NVIDIA system is even “DGX SuperPOD” :).
[1] https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Dat...
[2] https://devblogs.nvidia.com/dgx-superpod-world-record-superc...
The DGX-2h is a beast! Don’t parse this as “huh, TPU Pods are about the same as just a few V100s”. The data sheet [1] is probably the easiest to follow, but their writeup is more informative [2].
These are souped up V100s, with awesome networking, which is pretty similar in style to a TPU Pod. So I’d say that they’re both purpose built systems for distributed ML training. The name for the NVIDIA system is even “DGX SuperPOD” :).
[1] https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Dat...
[2] https://devblogs.nvidia.com/dgx-superpod-world-record-superc...
I think it’s probably because the benchmark isn’t optimized for TPU Pods. Check out the BERT in 76 minutes paper for how you need to rethink the training regime to take advantage of pods.
Yes, Cloud TPU Pods are designed to train much larger models on much larger datasets. And, as you mention, if you are willing to adjust your model architectures and training algorithms to take full advantage of the hardware, you can sometimes achieve substantial gains.
Author of the blog post here.
As mentioned in other comments, I'd recommend doing a performance-per-dollar comparison in addition to looking at this pure performance comparison at maximum scale.
As mentioned in other comments, I'd recommend doing a performance-per-dollar comparison in addition to looking at this pure performance comparison at maximum scale.
I'm curious how (absolutely) efficient the Transformer training is even on the TPUs. The results from self-attention models are really impressive but unfortunately their topology makes them very difficult to implement efficiently in silicon. It often becomes purely a question of memory bandwidth, because you're not doing much math per weight on each iteration. I wonder if the speedup is from the use of on-chip HBM in the TPUs.
You mean this speedup?
> 1024 TPUs are twice as fast as 480 GPUs
Might it be because there are twice as many?
> 1024 TPUs are twice as fast as 480 GPUs
Might it be because there are twice as many?
Author of the blog post here.
We submitted multiple results using various Cloud TPU v3 Pod slice sizes to show the current achievable Transformer training efficiency at several scales:
https://mlperf.org/training-results-0-6
We're actively improving the whole TPU software stack, so training efficiency is likely to continue to increase over time.
We submitted multiple results using various Cloud TPU v3 Pod slice sizes to show the current achievable Transformer training efficiency at several scales:
https://mlperf.org/training-results-0-6
We're actively improving the whole TPU software stack, so training efficiency is likely to continue to increase over time.
What I'm really asking is: how much effective compute throughput are you able to get during Transformer training relative to the amount of theoretical raw compute available?
What's in the picture? Can't figure out the scale. Are those like server racks or breadboards or ...?
Here's a higher resolution version: https://storage.googleapis.com/gweb-cloudblog-publish/origin...
All you see is 8 server racks with colored network cables, switches and power supplies. The TPU ASICs themselves make up just a tiny part of this datacenter, you also have the printed circuit boards, cooling fins, 8x48 metal boxes, power and network cables, DC/DC or AC/DC converters at the bottom, fans or water tubes for cooling and airgaps.
My startup is trying to develop wafer scale integration where you collapse 2 racks of the network, metal boxes, power and cooling into a 300mm wafer immersed a 100 mm x 400mm box with a few fibers and three power cables coming out. That can save around 900% of the capital cost and orders of magnitude of power (especially if you put the box in a building where the waste heat is not wasted but used to heat water for showering and space heating).
As it is now, these datacenter customers and hyperscalers don't seem to care about the enormous waste and cost of paying for inefficient hardware. Considering the enormous cost and carbon emmission savings a wafer scale integration would bring (and the competitive advantage), you would be suprised how hard it is to get funding from them to develop it.
I suspect is also the main reason they don't care to publish normalized benchmarks for $/performance/joule, as it would demonstrate how wasteful it all is.
My startup is trying to develop wafer scale integration where you collapse 2 racks of the network, metal boxes, power and cooling into a 300mm wafer immersed a 100 mm x 400mm box with a few fibers and three power cables coming out. That can save around 900% of the capital cost and orders of magnitude of power (especially if you put the box in a building where the waste heat is not wasted but used to heat water for showering and space heating).
As it is now, these datacenter customers and hyperscalers don't seem to care about the enormous waste and cost of paying for inefficient hardware. Considering the enormous cost and carbon emmission savings a wafer scale integration would bring (and the competitive advantage), you would be suprised how hard it is to get funding from them to develop it.
I suspect is also the main reason they don't care to publish normalized benchmarks for $/performance/joule, as it would demonstrate how wasteful it all is.
> would be suprised how hard it is to get funding from them to develop it.
That's probably because there has been no evidence that wafer scale integration can actually work.
That's probably because there has been no evidence that wafer scale integration can actually work.
I agree there is no complete evidence of a full WSI yet. There are several recent papers on Wafer Scale Integration (WSI) and some WSI built (large sensors) with reasonable yields. There are many papers on partial problem solutions that, if combined in one project would yield a full working WSI with existing 7nm standard CMOS process. There are silicon interconnect fabric (SiIF) which is one step removed from a full WSI.
There are many commercial chips 1/70th the size of a WSI already and some unpublished WSI results.
Server racks. The colored things are all cables.
Seriously.
Transformer: 1024 TPUs are twice as fast as 480 GPUs.
Resnet50: 1536 GPUs are about as fast as 1024 TPUs.
SSD: 1024 TPUs are twice as fast as 240 GPUs.
Great.