This $23B number that gets thrown around is not the increase to the public. The wording in the referenced report is
> Based on actual auction clearing prices and quantities and uplift MW, inclusion of existing and forecast data center load growth resulted in a combined total increase in capacity market revenue for the 2025/2026 BRA, the 2026/2027 BRA, and the 2027/2028 BRA of $23,100,955,341.
This is the increase in revenue to PJM from adding datacenter customers, and includes both the amount that datacenters paid as well as the amount that other customers paid due to higher prices from datacenters. So Fortune calling it an increase to "the public" means that they didn't read the report they are using as their source and are probably just repeating what they thought someone else meant.
Bloomberg in the past worded it as "data centers will add at least $23 billion to customer bills" in April and "added a minimum of $23 billion to customer bills" in February. Which while technically correct (datacenters are customers) seems meant to be misleading. And now that's the number that's getting thrown around as the increase to "the public".
The part I don't get is that the journalists could just give the actual number for the quantity that they are referring to (the amount that non-datacenters paid due to higher rates due to datacenter loads): when I calculated it a few months ago I think it was something like $16 billion rather than $23 billion. I feel like the story would have the same impact if the headline number was $16B as $23B, but $16B has the benefit of not being a misrepresentation of the situation.
---
Also I would definitely recommend checking out the PJM BRA report. It's a bit dense but not too hard to follow, and my personal takeaway was that the PJM market is just very dysfunctional and they are blaming the datacenters instead. I thought SemiAnalysis had a good analysis of it: https://newsletter.semianalysis.com/p/are-ai-datacenters-inc...
Fwiw this got changed about a week ago, where they changed the logic to match the documentation rather than default to sending your prompts to their servers. This is why so many people have noticed this happening but if you ask an AI about it right now it will say this is not true.
Personally I think it's necessary to run opencode itself inside a sandbox, and if you do that you can see all of the rejected network calls it's trying to make even in local mode. I use srt and it was pretty straightforward to set up
I think it's interesting that they dropped the date from the API model name, and it's just called "claude-opus-4-6", vs the previous was "claude-opus-4-5-20251101". This isn't an alias like "claude-opus-4-5" was, it's the actual model name. I think this means they're comfortable with bumping the version number if they want to release a revision.
They are definitely capable of writing such statements, which you can see in their enterprise products. In my Google Workspace gemini app it says pretty prominently and clearly:
Your [ORGNAME] chats aren’t used to improve our models
So they definitely understand that people want to hear that their data isn't being used for training, and they know how to say it clearly and reassuringly. Which makes the omission of that in their consumer products more telling in my view.
says Google's datacenter water consumption in 2023 was 5.2 billion gallons, or ~14 million gallons a day. Microsoft was ~4.7, Facebook was 2.6, AWS didn't seem to disclose, Apple was 2.3. These numbers seem pulled from what the companies published.
The total for these companies was ~30 million gallons a day. Apply your best guesses as to what fraction of datacenter usage they are, what fraction of datacenter usage is AI, and what 2025 usage looks like compared to 2023. My guess is it's unlikely to come out to more than 120 million.
I didn't vet this that carefully so take the numbers with a grain of salt, but the rough comparison does seem to hold that Arizona golf courses are larger users of water.
Agricultural numbers are much higher, the California almond industry uses ~4000 million gallons of water a day.
I was also surprised when someone asked me about AI's water consumption because I had never heard of it being an issue. But a cursory search shows that datacenters use quite a bit more water than I realized, on the order of 1 liter of water per kWh of electricity. I see a lot of talk about how the hyperscalers are doing better than this and are trying to get to net-positive, but everything I saw was about quantifying and optimizing this number rather than debunking it as some sort of myth.
I find "1 liter per kWh" to be a bit hard to visualize, but when they talk about building a gigawatt datacenter, that's 278L/s. A typical showerhead is 0.16L/s. The Californian almond industry apparently uses roughly 200kL/s averaged over the entire year -- 278L/s is enough for about 4 square miles of almond orchards.
So it seems like a real thing but maybe not that drastic, especially since I think the hyperscaler numbers are better than this.
I've found a method that gives me a lot more clarity about a company's privacy policy:
1. Go to their enterprise site
2. See what privacy guarantees they advertise above the consumer product
3. Conclusion: those are things that you do not get in the consumer product
These companies do understand what privacy people want and how to write that in plain language, and they do that when they actually offer it (to their enterprise clients). You can diff this against what they say to their consumers to see where they are trying to find wiggle room ("finetuning" is not "training", "ever got free credits" means not-"is a paid account", etc)
For Code Assist, here's their enterprise-oriented page vs their consumer-oriented page:
I agree in general, but I think some important context here is that the author of this post was previously on the OpenAI board (the board that fired Sam Altman).
The worst part to me is the privacy nightmare with AI Studio. It's essentially impossible to tell whether any particular API call will end up being included in their training data since this depends on properties that are stored elsewhere and are not available to the developer -- even a simple property such as "does this account have billing enabled" is oddly difficult to evaluate, and I was told by their support that because I at one point had any free credits on my account that it was a trial account and not a billed account even though I had a credit card attached and was being charged. I don't know if this is true and there is no way for me to find out.
At some point they updated their privacy policy in regards to this, but instead of saying that this will cause them to train on your data, now the privacy policy says both that they will train on this data and that they will not train on this data, with no indication of which statement takes precedence over the other.
There are a few conditions that take precedence over having-billing-enabled and will cause AI Studio to train on your data. This is why I personally use Vertex
The benchmark numbers don't really mean anything -- Google says that Gemini 2.5 Pro has an AIME score of 86.7 which beats o3-mini's score of 86.5, but OpenAI's announcement post [1] said that o3-mini-high has a score of 87.3 which Gemini 2.5 would lose to. The chart says "All numbers are sourced from providers' self-reported numbers" but the only mention of o3-mini having a score of 86.5 I could find was from this other source [2]
It's "experimental", which means that it is not fully released. In particular, the "experimental" tag means that it is subject to a different privacy policy and that they reserve the right to train on your prompts.
2.0 Pro is also still "experimental" so I agree with GP that it's pretty odd that they are "releasing" the next version despite never having gotten to fully releasing the previous version.
I believe that at least in the past the entertainment industry would try to detect someone seeding a file before going after them. The idea being that someone downloading is receiving a copy (not illegal), and the act of making the copy (illegal) was done by the seeder. I'm not sure to what degree this was an established requirement vs them trying to avoid ambiguity, but my point is that this framing by Meta isn't novel. I'm not expressing a judgment on whether it's correct or if it's good.
I think people here might like Oliver Burkeman's books where he talks about this stuff a lot. I loved his book "Four Thousand Weeks", and there is a new follow-up "Meditation for Mortals" which I have not read yet but seems to be well-received.
He's one of the few people I've seen address what I think is the key difficulty with this sort of stuff: that you can think think that you're addressing procrastination/perfectionism when actually you are engaging in it (with a target of fixing your procrastination/perfectionism). It's a difficult situation to break out of, because it seems like any effort to break-out would necessarily have this sort of grasping, but I think he (and Buddhist meditation) talk a lot about that key challenge.
This reminds me about the Semantic Web, which was a movement explicitly about making the web more understandable to machines. I don't agree with the ideas and I think a lot of other people were also skeptical, but I bring it up to say that some people take the other side of your argument rather seriously and that there's a lot of existing debate on the topic. Here's Tim Berners-Lee talking about this way back in 1999:
> I have a dream for the Web [in which computers] become capable of analyzing all the data on the Web – the content, links, and transactions between people and computers. A "Semantic Web", which makes this possible, has yet to emerge, but when it does, the day-to-day mechanisms of trade, bureaucracy and our daily lives will be handled by machines talking to machines. The "intelligent agents" people have touted for ages will finally materialize.
I quoted this from https://en.wikipedia.org/wiki/Semantic_Web since the original reference was a book that is not openly accessible. Also I think it's funny that he's talking about agents in exactly the same way that people do now.
A while ago I talked with someone who was working on clang-format and they said they tried this (at Google, I think) and the results were not good: they found people write different code depending on the format. For example, code written to fit in 120 columns but then formatted to 80 columns will look worse than code written for 80 columns, due to minor variations in verbosity and variable names and what not.
I notice this myself a bit when I switch from a fullsize monitor to a laptop screen.
I think it's fascinating that when the topic is "shell companies" that the HN discourse is essentially "if they have nothing to hide then they don't need secrecy". I think that if the article were about linking "tor users" with their secret owners then we would see the opposite stance being taken.
I'm not taking a position here, and I'm not saying even that these stances are necessarily contradictory, but just that the blanket argument "X shouldn't get to be secret because I don't think they have a legitimate reason" doesn't differentiate between these two cases.
I switched to Kagi a month ago, and initially I was pretty skeptical because a lot of the excitement sounded kind of hype-y and anti-Google.
But actually Kagi is quite good and definitely worth it. I have regained the expectation that when I search for something I will find the thing that I want, and I hadn't realized how much I had lost that with Google. It's hard to demonstrate this because I think it's an accumulation of many small improvements, so I encourage people to give it a try and see for themselves.
I do worry that this won't last forever -- for example, I think the AI features are being provided below-cost to gain market share, and it does worry me that they're spending so much money on these tshirts. But I can always switch away later so I don't worry about it that much.
I thought an interesting point was the liquid cooling -- unclear how important this is to them, but I'm guessing it means that they designed it with a TDP that requires liquid cooling.
This (wanting higher density) is the opposite of the trade-off that I was expecting. In my (limited and out of date) experience, power was the limiting factor before space, and I believe AI racks have very high power draws already.
I would have guessed this would be because larger nodes would be better for AIs tight communication patterns, but they specifically call out datacenter space as the constraint. Curious if anyone knows more about this