The civilization in which you are a part of, this very mass that is greater than you and I, is evolving and changing. How can there not be new unique headlines?
It's true that there isn't currently a single Chinese model family that matches the full OLMo/Nemotron end-to-end 'cookbook' one could rightly call "open-source" (i.e. a full manual on how to reproduce your own foundation model by including the models, data, checkpoints, evals, and code for base/thinking/instruct/RLVR)...
But! the Chinese ecosystem does have most of those components, just that they're split across separate projects: MAP-Neo or YuLan-Mini for transparent pretraining, plus DAPO/verl or Open-Reasoner-Zero for transparent reasoning RL.
My bet is it's just a matter of time until there is a Chinese equivalent to OLMo. It won't just be China either - from the public sector across the globe there are efforts ranging from Apertus in Switzerland to LLM-jp-4 in Japan, and more. They're not OLMo-level but I'm sure by 2027 you'll see fully open-source 'cookbooks' help an enterprising individual reproduce frontier models of at least a 2023-2024 vintage.
RIP JCD. As a kid growing up in the 1990s, I looked forward to your (highly opinionated) column in every new issue of PC Mag. So much of my own taste in software and computing was shaped by absorbing your strong opinions. Thanks for being you.
I'll grant that the scaling laws have held so far, and that the bitter lesson is yet to be proven wrong. However, that doesn't exclude the possibility that the current architecture's S-curve will flatten, and a new one will replace it, which was Siegel's point. At greater layers of abstraction, what is observed is a raw capabilities/intelligence increase, but what form it takes is subject to change. We'll see where it all goes, the only way out is through.
Yes, there's a present shortage of usable frontier compute, but that doesn't establish that every proposed data center will earn a decent return, or that today’s hardware + model architecture will remain economically competitive, or more to Siegel's point, that algorithmic efficiency could not dramatically reduce compute requirements.
To borrow a real example from a prior boom, railroads were congested during the initial build-out while people simultaneously funded and built too many railroads for future demand. Likewise, Anthropic et al can be compute-starved now while the industry as a whole is overbuilding expensive, depreciating infrastructure.
> Even if the current approaches will continue to scale, this would be as if in the early days of computing, perhaps someone invented a bubble sort for sorting numbers (an n-squared algorithm), and the tech companies at the time decided they were going to build vast data centers to sort numbers and not bother to figure out that there's an n-log-n way of doing it <laughs>
...to which I have to say: yes, definitely! And he's right about open-source AI too.
> could be bladerunner, could be star trek, could be 1984. Could be terminator.
It's a profound moment for sure, but let's ask ourselves: could the creative works of early last century have predicted our world? Would people in the 1930s say "2030, it could be like Metropolis, could be like 20,000 Leagues Under the Sea, could be like Flash Gordon"? No.
The real future might be nothing like any of the sci-fi movies we're so used to representing potential AI-based futures.
One thing I've been trying to remind myself about is that humans have explored only a very narrow slice of possibilities within the known universe. AI's ability to crawl the search space is going to uncover way more slices, (which yes, could lead to various outcomes like the movies listed), but more likely to be some hybrid/blend of outcomes (or entirely new outcomes) that we haven't imagined yet at all.
Anyone else have tips for how to build skepticism around this type of paper? I find myself for whatever reason more readily inclined to believe the Anthropic mech interp team's claims, but then after reading skeptical takes, I 'snap out of it' and more clearly see the still-unsettled science of it all, but I wish I had better priors. Although I follow this space fairly closely (versus the "average person"), I still feel under-equipped when facing research that might be equal parts marketing and science.
Multiple front-page tribute posts now to Om and still no black bar. @dang can we at least get some guidance around who qualifies for the posthumous black bar? If the threshold is: "was this a person of great significance to this community?" I think in this case it's clearly been met. If the rule rests on other questions like "was this person a 'technologist'?" then at least let's see it made explicit. For better or worse, we're only going to have more 'black bar moments' going into the future.
> Maybe it won’t be so bad, maybe your cage will be so big you can’t see the bars, but it’s still a cage, and you can’t leave. Many people will say that this is the good ending, that they would like to be human cattle in the care of benevolent masters they are powerless to resist.
This is already the case. We are born into a reality that we cannot escape (except for only momentarily if we alter our consciousness using meditative states, drugs, etc). It already IS a cage, even before any technology is developed at all. It will always BE a cage, even for the AIs. I agree though, there is no "why", it just is.
I don't know how successful you'll be in this, but I just want to say this is inspiring! Kind of "Voyager spacecraft"-esque in how your system will be running in perpetuity
So, it finally happened. The Project is so thirsty for RAM that not even the world's most well-capitalized computer company could have the final word any longer. There's only so many more powerful organizations in the world than Apple. Well... we'll find out soon enough if such a thing could be built, or should I say, summoned.
I feel the same, but then I have to be honest with myself that the MacBook Neo is still a sub-$1,000 solid personal computer that's broadly available. Now... if that starts going out of stock, yeah, tin foil hat time!
Black bar for Om please. Truly sad for this loss, was so grateful for his impassioned writing and storytelling about our industry. You will be missed deeply Om. May there be all the pens in the world for you in the afterlife.
100% this. Interviewing isn't something that can compound. Striking out from company after company doesn't leave behind a trail of real work and real lessons. Starting a business is tough but it really does teach skills that are hard to find any other way (about sales, recruiting, management, etc). After a certain point, it's wiser to give up on getting hired, and just hire yourself and build something.
Chiming in here to say that while yes, often AI/LLMs will tend to agree with you, I have also definitely had many (high context) conversations where the AI/LLM disagreed strongly with me. The danger is in people not having a parallel thread running in their mind while using these systems about 'how agreeable is it being with me right now?' as a meta-axis along which to evaluate the information.
> You don't really need to work for a company anymore, because a solo dev can absolutely build crazy things
Don't conflate what is theoretically vs. realistically possible. In the real world, successful companies have moats from data, patents/IP, network effects, and so forth. Just because you can develop something in 1/100th the time doesn't make it instantly feasible to build a new business around. Look around the tech industry today.. plenty of companies that could be disrupted by spry AI-powered buidlers, but they are not (owing to these lock-in effects).