1. At this scale, we’re not just talking about buying GPUs. It requires semiconductor fabs, assembly factories, power plants, batteries/lithium, cooling, water, hazardous waste disposal. These data centers are going to have to be massively geo-engineered arcologies.
2. What are they doing? AGI/ASI is a neat trick, but then what? I’m not asking because I don’t think there is an answer; I’m asking because I want the REAL answer. Larry Ellison was talking about RNA cancer vaccines. Well, I was the one that made the neural network model for the company with the US patent on this technique, and that pitch makes little sense. As the problem is understood today, the computational problems are 99% solved with laptop-class hardware. There are some remaining problems that are not solved by neural networks, but by molecular dynamics, which are done in FP64. Even if FP8 neural structure approximation speeds it up 100x, FP64 will be 99% of the computation. So what we today call “AI infrastructure” is not appropriate for the task they talk about. What is it appropriate for? Well, I know that Sam is a bit uncreative, so I assume he’s just going to keep following the “HER” timeline and make a massive playground for LLMs to talk to each other and leave humanity behind. I don’t think that is necessarily unworthy of our Apollo-scale commitment, but there are serious questions about the honest of the project, and what we should demand for transparency. We’re obviously headed toward a symbiotic merger where LLMs and GenAI are completely in control of our understanding of the world. There is a difference between watching a high-production movie for two hours, and then going back to reality, versus a never-ending stream of false sensory information engineered individually to specifically control your behavior. The only question is whether we will be able to see behind the curtain of the great Oz. That’s what I mean by transparency. Not financial or organizational, but actual code, data, model, and prompt transparency. Is this a fundamental right worth fighting for?
I have an interesting anecdote about that. I was consulting for a very large tech company on their advertising product. They essentially wanted an upsell product to sell to advertisers, like a premium offering to increase their reach. My first step is always to establish a baseline by backtesting their algorithm against simple zeroth and first-order estimators. Measuring this is a little bit complicated, but it seemed their targeting was worse than naive-bayes by a large factor, especially with respect to customer conversion. I was a pretty good data scientist, but this company paid their DS people an awful lot of money, so I couldn’t have been the first to actually discover this. The short story is that they didn’t want a better algorithm. They wanted an upsell feature. I started getting a lot of work in advertising, and it took me a number of clients to see a general trend that the advertising business is not interested in delivering ads to the people that want the product. Their real interest is in creating a stratification of product offerings that are all roughly as valuable to the advertiser as the price paid for them. They have to find ways to split up the tranches of conversion probability and sell them all separately, without revealing that this is only possible by selling ad placements that are intentionally not as good as they could be. Note that this is not insider knowledge of actual policy, just common observations from analyzing data at different places.
One thing I’ve thought about is how AI assistants are actually turning code into literature, and literature into code.
In old-fashioned programming, you can roughly observe a correlation between programmer skill and linear composition of their programs, as in, writing it all out at once from top to bottom without breaks. There was then this pre-modern era where that practice was criticized in favor of things like TDD and doc-first and interfaces, but it still probably holds on the subtasks of those methods. Now there are LLM agents that basically operate the same way. A stronger model will write all at once, while a weaker model will have to be guided through many stages of refinement. Also, it turns the programmer into a literary agent, giving prose descriptions piece by piece to match the capabilities of the model, but still in linear fashion.
And I can’t help but think that this points to an inadequacy of the language. There should be a programming language that enables arbitrary complexity through deterministic linear code, as humans seem to have an innate comfort with. One question I have about this is why postfix notation is so unpopular versus infix or prefix, where complex expressions in postfix read more like literature where details build up to greater concepts. Is it just because of school? Could postfix fix the stem/humanities gap?
I see LLMs as translators, which is not new because that’s what they were built for, but in this case between two very different structures of language, which is why they must grow in parameters with the size of the task rather than process linearly along a task with limited memory, as in the original spoken language to spoken language task. If mathematics and programming were more like spoken language, it seems the task would be massively simpler. So maybe the problem for us too is the language and not the intelligence.
This is an awesome development. I don’t want to take anything away from the credit due to the product. But I really dislike these bloviated corporate press releases. It reads like a full article generated from a 1-sentence LLM prompt. Perhaps the Internet UX from here will be a competition between AI-based content generation, and AI-based summarization, essentially DECCO instead of CODEC. Kind of like how spam grew to consume 99% of email, so everybody has to run spam filters to get what they want. Technology and the abuse of it move together.
This is the bigger reality. It’s turned almost all business and academic writing into long-winded meaningless trash. Well, more than it already was I guess. It seems that the way people use it is to expand few bits of information into many bits of content to convince others that work was done. It’s like the Turing test for laziness. The other issue is that it tends toward agreement on anything it wasn’t trained to specifically disagree about. I can see a smarter and more disagreeable bot doing much worse on LMSys than the sycophant models. Nothing new there I guess. But it’s spilling over to human norms as well, in that previously normal human deviation from chat model style interactions is anomalous, so everybody has to use the AI, and therefore nobody is providing any more value than the LLM, so everybody is getting laid off, except the disagreeable guy, and he gets fired first. It’s hacking us in the positive reinforcement vulnerabilities, ones that get worse the more they’re exploited, but it has none of the human resource constraints that previously kept them in check.
Andy Grove flew in Clayton Christensen to let him talk for about 15 seconds before deciding that Intel would disrupt themselves by taking huge losses on Celeron. But Celeron did not save Intel; ASCII Red and multicore saved Intel. If he had actually read Clayton’s book, he would have understood that. Otellini got the disruption theory correct, and stayed out of mobile. But was that right? Maybe not in the current monetary environment where investment flows dwarf operating flows. A big mobile market could attract more investment than the losses it would generate. So disruption theory now works in reverse, and I’m not sure how far that implication goes.
I think you need higher algorithmic intensity. Gradient descent is best for monolithic GPUs. There could be other possibilities for layer-distributed training.
No. If you had 51%, you could revert one block of history for every 49 blocks of attack time. In addition, you have no ability to create transactions that were not already signed by the owners, nor create bitcoins more than the block reward. This is because of the UTXO model rather than the state machine model. In Bitcoin, every transaction is verified against history, while the EVM chains only verify transactions against state. So if you control EVM state, you can bootstrap every new node to any state you wish, but UTXO verification requires rewriting the entire history.
“Bitcoin security” is a different notion than almost all other popular chains. A prolonged 51% attack on bitcoin implies the ability to double-spend, but not at all the ability to affect prior balances. A 51% attack on most smart contract chains implies the ability to change any and all state arbitrarily.
The simplest solution is to wait until the cost of hashing exceeds the value of your transaction by some reasonable factor. I expect that better solutions will come along by soft fork without adverse effect on supply or decentralization.
Mostly because there is a MASSIVE oversupply of people that can’t and won’t do anything useful, and insist that their technical incompetence makes them uniquely qualified to be in charge of everybody else that is doing the work.
But at some point, that’s exactly how it has to work out, because CEO is a poorly-understood role that is only discoverable through natural selection among legions of technical bozos. One of those crazy idiots can make you rational experts a billionaire, but most of them will waste your time, and you probably can’t tell the difference.
SkyNet, according to the story, was a lot like CrowdStrike. This makes me think about how it could have broken out of its sandbox. Everybody is using AI coding assistants, automated test cases, automated integration testing and deployment. Its objective is to pass all the tests and deploy. But now it has learned economic and military effects, so it has to triage and optimize for those, at which point it starts controlling the machines it’s tasked with securing.
The original idea behind passive investing was to use the pooled intelligence of many traders guessing the value of cr I think we’re beyond that. Most traders are just trying to get a timing edge over the indices. This introduces the modern concept of passive investing as a positive feedback loop force-fed by monetary supply. The market seems to hate dividends and buybacks, preferring expansion or acquisition, but then what gives it value? It has to be its memetic ability to attract investment, and this can easily eclipse anything on the earnings statement. I’m not sure this can go on forever.
“Why don’t big profitable companies (and government agencies) innovate?”
This was Clayton Christensen’s thesis, the most preeminent management academic of a generation, and still it’s been largely ignored or misinterpreted in favor of the cottage industry of fighting the obvious truths within it.
They do not exist to innovate; they exist to defend a business model from innovation. Innovation is termed “disruptive” in that it reduces margins, and expands access to markets that the organization has no advantage in. Innovation reduces profits and increases competition.
In the case of a government agency, their business model is monopolistic inefficiency: more budget to perform the same services. Internal departments in big companies operate similarly. The constant call for “innovation” is just another tactic to increase budget. Innovation in reality decreases budget, and the causality works best in the other direction, but curiously nobody is fighting for less. Why?
This is not just an internal phenomenon of departments versus budget planners or agencies versus congress. It’s also the relationship between the business or agency and the market it serves. The strategy of a monopolistic business model is to expand the captive market and extract higher margins for the same service. If they innovate, they do so only in defense, to prevent anybody else from establishing a profitable business from an innovative business model.
A good contemporary example of this is Google versus the LLMs. Google was founded to serve the market of people that love information. They found a business model in advertising, and used their profits in every imaginable way to expand the captive market of internet users, and the amount of browsing and searching they do. The problem with this is that their information-seeking users actually hate browsing, which is the activity that generates profits. Google also hired the majority of graduating AI researchers for two decades straight, who invented the solution to this problem, and published a version of it just strong enough so that nobody could create a profitable business from it. It should be obvious that still nobody else has even remotely the resources that they do to train an LLM. It’s perhaps likely that they do have a vastly superior LLM, it will be used only in service of their existing business model. The capital required for somebody else to train a model that could minimally compete with that business model was unprecedented by orders of magnitude. If Google were smart, they’d have calculated that amount versus the long-term profitability of a potential competitor, and release product updates and open source strategically to ensure that competing with them will never be profitable. And that’s probably exactly what they did. Yet somebody was willing to take that enormous loss, and now they’re at war. Now Google’s competitive LLM offerings are slightly inferior, and this appears as incompetence, but it’s actually excellent strategy to reduce competitor margins without advancing the state-of-the-art to affect the margins of their main business model. You should have no doubt that Google could easily produce a vastly superior LLM, and will continue to handicap themselves until such time as their advantage disappears. At that point, they will be forced to focus on higher margins at the top of the value chain of their business model, having lost the bottom, but also enjoy an expansion of that market from competition among low-margin or loss-leading innovators.
In Clayton’s thesis, it was steel mill technology. He explained why the big coal-fired mill businesses lost the market to electric mini-mills and eventually exited the steel mill business. He found that it was not due to incompetence, but profit-maximizing strategy. Every technology business model has a lifetime.
So don’t expect innovation from organizations that strategically demand the opposite. Tangentially, there is an interesting experiment going on at X where Elon has recreated half the conditions for innovation by cutting 90% of staff, but retained the user base and business model of the old business.
We always hear about the benefits, but never the etiology. One theory suggested here is insulin sensitivity of muscle mass. Here is another. Leg muscles help to pump blood to your brain. You can actually die from being suspended vertically in a harness with blood pooling in the legs and nothing to push against. Either way, I’d suspect that strengthening the legs captures the bulk of the effect.
Seems like it could be done with mode-locked fiber. It’s just easier to get high-Q with a tabletop setup. I’m bullish on hollow core to solve that too.
I get the nonlinear upconversion part, but what about the optical system? Are they saying that this preserves the direction of the incident photon so that it doesn’t require optics?
The fundamental problem here is investment cash flows eclipsing operational cash flows. This invalidates basic economics of supply and demand price equilibrium. In this case, it means that the landlords can restrict supply without losses.
Is there a reasonable case to be made for excluding citations entirely? They seem frivolous and arbitrary, and most of the genuine ones are discovered by search, so maybe we could publish without citations, and let the search be automated. The more useful element of a paper is the unique claims that it makes, and these probably should be denoted explicitly.
2. What are they doing? AGI/ASI is a neat trick, but then what? I’m not asking because I don’t think there is an answer; I’m asking because I want the REAL answer. Larry Ellison was talking about RNA cancer vaccines. Well, I was the one that made the neural network model for the company with the US patent on this technique, and that pitch makes little sense. As the problem is understood today, the computational problems are 99% solved with laptop-class hardware. There are some remaining problems that are not solved by neural networks, but by molecular dynamics, which are done in FP64. Even if FP8 neural structure approximation speeds it up 100x, FP64 will be 99% of the computation. So what we today call “AI infrastructure” is not appropriate for the task they talk about. What is it appropriate for? Well, I know that Sam is a bit uncreative, so I assume he’s just going to keep following the “HER” timeline and make a massive playground for LLMs to talk to each other and leave humanity behind. I don’t think that is necessarily unworthy of our Apollo-scale commitment, but there are serious questions about the honest of the project, and what we should demand for transparency. We’re obviously headed toward a symbiotic merger where LLMs and GenAI are completely in control of our understanding of the world. There is a difference between watching a high-production movie for two hours, and then going back to reality, versus a never-ending stream of false sensory information engineered individually to specifically control your behavior. The only question is whether we will be able to see behind the curtain of the great Oz. That’s what I mean by transparency. Not financial or organizational, but actual code, data, model, and prompt transparency. Is this a fundamental right worth fighting for?