In case anyone misses the links, this is twinned with two other superb posts - one about general lessons the author learned over the course of the project
I'm a huge fan of using simulations to ground qualitative arguments. While the sims usually need to be fine-tuned so extensively as to leave them open to claims of 'overfitting', the benefit is that it nails the assumptions in your argument to the church door.
Abstract:
Patterns of political unification and fragmentation have crucial implications for comparative economic development. Diamond (1997) famously argued that “fractured land” was responsible for China's tendency toward political unification and Europe's protracted political fragmentation. We build a dynamic model with granular geographical information in terms of topographical features and the location of productive agricultural land to quantitatively gauge the effects of “fractured land” on state formation in Eurasia. We find that either topography or productive land alone is sufficient to account for China's recurring political unification and Europe's persistent political fragmentation. The existence of a core region of high land productivity in Northern China plays a central role in our simulations. We discuss how our results map into observed historical outcomes and assess how robust our findings are.
Did a double-take seeing Maldacena's name on this. He's better known for discovering AdS/CFT, which is the foundation of a lot of modern work on quantum gravity.
If it's 100x from increased investment and 10x from short-term efficiency gains, yeah you'd expect $4/page. Model compression or some other tech might make it more efficient in the long-run.
Training an AI in 2020 is best thought of as a capital investment. Like digging a mine or building a wind farm, the initial investment is very large but the operating costs are much lower, and in the long run you expect to get a lot more money out - a lot more value out - than you put in.
Training GPT-3 cost $5m; running it costs .04c per page of output.
When speaking of billion dollar investments, a billion dollar industry is not substantial. Google and Facebook's industries are advertising, at $600bn/year. Amazon's industry is retail, at $25tn/year.
What's opened up by the GPT-3 and its prompt-programming abilities is services, without qualification. That's $50tn/year, and capturing some tiny percentage of it is what's needed to make a billion-dollar investment worthwhile.
That said, I admit this isn't the mindset most people take when they read 'substantial'.
e: I changed the wording from 'substantial' to 'transformative', thanks!
There's no shortage of cities where central space has become cheaper after their industry has collapsed, and they do not have a great record of turning vibrant. The rust belt and the north of England are the first areas to mind.
For anyone else just looking for an outline of the new algorithm: last two paragraphs of p4, first two paragraphs of p5.
Coming from a position of a few abstract algebra classes in college many years ago, all the words and notation are familiar but I am a long way from being able to follow it.
More generally, the best commentary I've seen on GPT-3 comes from following @gwern and keeping an eye out for whoever else they frequently interact with.
Most other commentary I've seen either goes too far into reactionary skepticism, too far into to-the-moon style hype, or just straight-up gets stuff wrong by not being aware of some technical detail (importance of prompt design, BPE limits).
https://clemenswinter.com/2021/03/24/my-reinforcement-learni...
and one history of the project
https://clemenswinter.com/2021/03/24/conjuring-a-codecraft-m...