Yes, others should be able to do this too, and the change is larger memory chips.
The next big step is Medusa Halo, which will have a 384-bit LPDDR6 interface. Those should be able to support 256GB at release, with 512GB coming later with larger chips (But I don't know to what extent people should trust the memory vendor roadmaps.) I'm not sure if they will be out in a year. Probably will be in a year and half.
Do not base products on models that are not open-weights. Doing it is like building a product on someone else's platform, you are entirely at their mercy, and even when they don't have any reason to hurt you, you are tiny enough that if any policy they want to enact hurts you as a side effect, no-one is going to care.
You don't have to self-host the open-weights model, you just need to be able to source it from multiple providers.
Using the closed vendor models maybe made sense when open-weight models lagged so far behind, but that time is now gone.
Not because rom is better than flash, but because the critical part is distributing the memory with the compute. Mask ROM is just the densest way of embedding memory with logic. Instead of having a large pool of memory connected to the separate execution units with a bus, each execution unit locally has the rom that it uses. Data movement is >90% of energy use in modern ai accelerators, removing it as far as practical is how they get performance and silicon and energy efficiency.
Taalas only uses SRAM for the KV cache and the activations, the weights are in mask rom in the metal layers.
If they designed this right, it means that once they have a model, so long as they keep the hyperparameters fixed they can change the weights much faster than it takes to spin up a completely new chip, essentially at a cost of doing a minor revision.
It directly impacts performance when playing twitch fps games, any input lag causes measurable difference in score. This is before people even notice it happening.
No. The reason was not economic, the reason was that the government polled the people living there and found that support for remaining British is ~100%.
Oil was found later, the fisheries were never worth maintaining the island for.
The Australian grid presently curtails ~7-18% of production every single day between 11:00 and 14:00.
I believe that incentivizing people to acquire batteries is precisely the purpose of the policy. It's good for the grid for there to be a lot of storage at the edges. As I understand it, the 24kWh cap is subject to annual review, with it being reduced/the policy being soft phased out once curtailment is no longer necessary.
The problem is, we don't just have an absence of evidence, we have evidence of absence. The area has been widely excavated, and there is a clear continuity of settlement with the same pottery, culture, and religion. There is simply no trace of any large-scale population movement. As far as we can tell, the same people continued living in the area in the same way, worshipping the same gods (still plural for way longer) with the only large change being the yoke of the nearby great powers going away with the collapse.
This of course doesn't mean that there cannot be a trace of truth in the story! It just has to have been morphed substantially over time. For example, it was common in the time to kidnap and move foreign nobility and artisans, while no-one much cared about the identity of the average farmer or goatsherd. It could well be that "the people of Israel" who were kidnapped meant the people who actually mattered, ie, a fairly small upper class group, who could move from the Nile valley to the levant without leaving much trace in either society.
I'm still personally partial to the observation that the story seems to originate during the Babylonian captivity, and the situation of the story greatly mirrors the conditions they were living under, but while complaining about their Babylonian overlords was probably not allowed, writing stories about the plucky underdogs outwitting the horrible Egyptian overlords with divine assistance was fine, even if it contained themes of returning home and of liberation from foreign rule. (Note that Egypt was the main rival of Babylon in this period, and the Kingdom of Judah was on-again off-again vassal of the Egyptians. The captivity was party imposed to prevent this relationship from continuing.)
It's not quite as bad as the parent made it out to be, the largest I've seen is 32kB per token (where sometimes, a token represents a byte, but usually it represents more than one.)
It's forced by the nature of how LLMs use vector embeddings for language.
Basically, a single token in a LLM is represented as a n-element vector, where n is the "hidden dimension", also known as model dimension. In order for the model to be smart, the hidden dimension needs to be large, on the order of 2^16 on top-tier models. Elements of this vector are typically quantized to 2-byte floats, or sometimes smaller. Every possible fact is embedded as a direction in this very many dimensional vector space, and a token is related to a fact if the vector representing that token points into a similar direction as that fact. You can do vector math about these things, famously for most trained models, if you find the vector embedding for king, man, woman and queen, and calculate king - man + woman, the result is very close to queen.
(Does that mean that there are 2^16 possible different kinds facts about things in this model? No, because high-dimensional geometry is very unintuitively powerful. The facts are not axis-aligned, and they don't need to be perfectly non-orthogonal. This matters, because the numbers of individual vectors you can fit into a single 2^16 dimensional space that are orthogonal with each other (all angles 90degrees) is of course 2^16. But, if you allow for almost orthogonal vectors, the number is larger than the amount of atoms in the universe. If this sounds wacky, for people with a CS background it can help to think it working a bit like a bloom filter, in that collisions are possible. Although in actuality they are theoretical, because 2^16 is a very large number.)
That's not how modern ethernet works at all. A single NIC talking directly to an another one has no collisions ever. Depending on what your channel is, either you have separate wires for the directions, or you are using a hybrid circuit (as in telegraphs, the term is so overloaded it's hard to google). Either way, packets going in one direction never wait for packets going in the other.
This is an example of common knowledge that is wrong. People look at their cash burn, assume that they spend this to subsidize inference, and get bonkers answers. Inference is not their largest expense.
Inference is cheap. Anthropic is only drastically subsidizing their plans if you count their training expenses as part of their costs.
Mostly training. Claude didn't just get to be so good at coding by magic, it was suddenly so good because they did truly staggering amounts of RLHF and RLAIF on it. They are still doing that today, on any tasks they can figure out how to evaluate it on. This is capex for them.
Their margins on inference are >90% today for tokens they sell (plans are hard to count, but still profitable). Based on what we know of it's size and architecture, running Opus is not more than 2x more expensive than running Deepseek v4 pro, for which tokens are available at under 10% of the cost of Opus. Again, the reason their margins are 50% is because they are spending so much on things that are not inference, not because inference is expensive.
> The cheap model providers have a much better chance of achieving that.
Anthropic can do it with a push of a button, once they calculate that it will provide them better profit than current pricing.
> What worries me about this is that Anthropic and OpenAI seem to have backed themselves into a corner of high costs. Can they reasonably decrease their prices by 20-50x to compete with DeepSeek or Xiaomi’s Mimo?
They have high prices, not high costs. They will obviously keep prices as high as they can for as long as they can, while keeping demand up. Once demand starts to fall, so will the prices.
> Are these models cheap because they are open weight and having hundreds or people stress test running them on different hardware helped to lower the cost? Or is it that they are being provided as loss leaders to drive the prices down?
Neither. They are cheap because they have neither technical edge nor brand power to keep the prices high, and so have to ask commodity prices for them.
People somehow still don't get it, despite everyone who studies the economics of it telling them: Inference is dirt cheap. Training is expensive, inference is cheap, and getting cheaper.
It won't make sense to run them after two years. The vendors will be limited on datacenter space, power and cooling, and there will be new hardware available that will run the same models at a fraction of the power.
A100 -> H100 was >3x tokens per joule, H100 -> B200 >10x. There are significant low-hanging fruit still available in architectural efficiency, and the vendors are chasing them.
This is the big risk for AI companies that I feel is not being sufficiently priced in. Almost none of the investments they are making are durable, the depreciation schedules for everything but the real estate should be less than 24 months. Until the hardware is stable enough that you only get double-digit % improvements per generation, it should almost be counted as opex.
All of them. And it's not just an os thing, it's usually printed in the bottom right corner of the M key, the same way € is in the corner of the E key.
The next big step is Medusa Halo, which will have a 384-bit LPDDR6 interface. Those should be able to support 256GB at release, with 512GB coming later with larger chips (But I don't know to what extent people should trust the memory vendor roadmaps.) I'm not sure if they will be out in a year. Probably will be in a year and half.