I don't mean to be shady, but there are plenty of details that they did release that show that they don't know what they're doing.
They make comparisons to FlashAttention-2 when FlashAttention-4 has been out (even if they wanted to stick to Hopper class GPUs for whatever reason there's still FlashAttention-3). The two orders of magnitude claim look like they're for prefill not next-token decoding, which is a bit duplicitous. Long context extrapolation experiments typically go well beyond 2x context length. Etc etc etc.
I never said they should have a full public disclosure, but I do think sharing something of substance helps build trust and also get people excited.
Lastly, frontier labs have other incentives than to eek out every dollar and cent. Having the most capable models, not the most cost effective, is of significantly higher priority as OpenAI and Anthropic march towards IPOs. The same is not necessarily true for Google/DeepMind, and one can see from their public releases alone for some of their open weight models that this may be more of a priority for them today.
Ahh cf my comment above. The cost of failure at scale is too high for a major to just take a new architecture/mechanism and implement it, especially because a) most claims papers make aren't rigorously tested and b) plenty of things that work at one scale do not work at the scale on which the labs operate. If they want to get acquired, then they should show that they know what they're doing. Otherwise, it looks sketchy.
I don't think it makes sense from a business perspective to hold off on details as a new lab. OpenAI will not implement new architectural changes unless they've tested the changes themselves internally. Even if someone claims some great innovation, they'd need to do scaling experiments to somewhere between the size of GPT-4 to GPT-5 before they'd decide it is worth it to implement themselves. Plenty of mechanisms that seem to work at one scale do not translate to the next.
Because the cost to OpenAI to make an architectural shift is far greater than the cost to a new lab to try something different, providing details is usually a net benefit for recruiting, building trust, getting acquired, etc. The lack of details is a poor business decision because it makes them seem untrustworthy.
I'm not advocating that they should open source their model, but there is already so much noise in the space and many bad papers that being cagey is a poor strategy for winning over talent, developers, etc.
I don’t understand why this lab is allergic to providing details on what they actually made, especially when Chinese labs are more than willing to share architectural specs/code/kernels (eg NSA/FSA, RAMBa, HISA, DSA LightningIndexer, etc). I don’t doubt that they’ve done something here, but the lack of details makes me default not trust this, particularly when this is the second time that they’ve released a “technical report” that just waxes poetic about the concept.
The article does a great job of highlighting the core disconnect in the LLM API economy: linear pricing for a service with non-linear, quadratic compute costs. The traffic analogy is an excellent framing.
One addition: the O(n^2) compute cost is most acute during the one-time prefill of the input prompt. I think the real bottleneck, however, is the KV cache during the decode phase.
For each new token generated, the model must access the intermediate state of all previous tokens. This state is held in the KV Cache, which grows linearly with sequence length and consumes an enormous amount of expensive GPU VRAM. The speed of generating a response is therefore more limited by memory bandwidth.
Viewed this way, Google's 2x price hike on input tokens is probably related to the KV Cache, which supports the article’s “workload shape” hypothesis. A long input prompt creates a huge memory footprint that must be held for the entire generation, even if the output is short.
Does anyone know how AI coding fits in with S174? If a person’s “coding” part of the job is primarily running prompts and checking code outputs (quality control and minor reprompting) with the remainder of the time used for other activities, does this count as software engineering?
It seems like an inevitable outcome of this is elaborate system-gaming to mitigate how much employees fall under S174…
I think the most interesting thing to me is they have multi-hop search & query refinement built in based on prior context/searches. I'm curious how well this works.
I've built a lot of LLM applications with web browsing in it. Allow/block lists are easy to implement with most web search APIs, but multi-hop gets really hairy (and expensive) to do well because it usually requires context from the URLs themselves.
The thing I'm still not seeing here that makes LLM web browsing particularly difficult is the mismatch between search result relevance vs LLM relevance. Getting a diverse list of links is great when searching Google because there is less context per query, but what I really need from an out-of-the-box LLM web browsing API is reranking based on the richer context provided by a message thread/prompt.
For example, writing an article about the side effects of Accutane should err on the side of pulling in research articles first for higher quality information and not blog posts.
It's possible to do this reranking decently well with LLMs (I do it in my "agents" that I've written), but I haven't seen this highlighted from anyone thus far, including in this announcement.
Higher weight polyphenols tend to taste less bitter than lower weight ones. Its more correlation than causation because I don’t think we precisely know why this is.
Yeah! Green tea gets fried or steamed right away to halt oxidation. That kills off some of the undesirable bitterness that masks some flavors that are even present in fresh leaves. It is not that Green doesn’t have any taste; it is that there are more guardrails over what flavors can appear and how distinct they can be.
It’s hard to build a product that shows you how clothes fit on someone like you as a B2B service. Retailers don’t want to showcase their clothes on anyone who isn’t anatomically perfect. Plus, if you try to source a more diverse set of “models” from real people wearing clothes, you run into the problem that most people are uncomfortable sharing photos of themselves “modeling” clothes publicly.
You’re also right that fit is only a part of the picture, and even the terms fit and style don’t quite capture what’s really going on you really want to see what clothes are going to look like on someone who looks like you and dresses like you (same preferences for fit, style, etc). Again, hard as B2B for sure.
I’ve been working on a B2C solution in this space for a while (fitfirst.app)…all too familiar with the nuances and intricacies in this space.
Fun fact and totally tangential: if you have two clothes with the exact same measurements and material that are dyed different colors, the darker dyed version tends to feel tighter than the lighter dyed one. Has to do with how the dye feels on the skin. It’s a nuance you can’t get from a photo or rendering.
My partner and I have been creating a database of over a thousand pant measurements that we've personally gathered over the past year or so. We've found it irritating that the only way to shop for apparel is pretty much by trial and error (fitting rooms, manually looking at size charts and hoping that they're accurate). I have a super small waist-to-hip ratio (and a super small waist) for a man, so I run into two problems
1. Usually stores don't carry the right waist size for pants I want to buy, so I can't try them on.
2. When I buy online, stuff that's in my size is usually too tight around the seat of my pants.
That's why I thought it'd be fun to create a way of browsing pants that let you see the differences between two pairs of pants so you could figure out if something is even worth trying to buy.
Over the past couple days, my partner and I put together a tool in D3 that overlays pants and compares their measurements everywhere.
Let me know what you think, and let me know if you have any technical questions about it.
I’ve started picking the bottom result of google on page ten for fun...you get some really wacky content that clearly isn’t optimized for SEO or isn’t relevant at all.
Example: last night I searched “Robin Williams bipolar” and got a page ten result that was a conspiracy theory on Taylor Swift being a psychopath.
Great question! The term skinny jean has changed over time.
What I could have been more careful in defining is what a "modern skinny jean" is: it is designed to be form fitting throughout the seat and leg. That means a tapered leg opening: the knee is larger than the hem. That also means the garment measurements are smaller than your own body measurements everywhere, except possibly at the leg opening.
Lycra/spandex/elastane had been incorporated in jeans in the 60s (I checked the Sear's catalogue to verify material composition when I was researching this), but the jean silhouette of the pants they sold were closer to what we'd call a "slim fit" today. Same with denim in the early 1970's: it hugged the body, but wasn't skin tight everywhere, and often was boot cut at the leg opening.
Fiorucci was the first to add stretch to denim with the purpose of making the jeans skin tight and tapered at the leg.
Hey everybody. A little bit about this. I've been measuring hundreds (looks like over a thousand by now) jeans with my partner over the past year. We've become a bit obsessed with turning clothes shopping into more of a science, mostly because we've found clothes shopping incredibly opaque and frustrating.
I've taken a bunch of the measurements and other data I've collected and turned it into a fairly comprehensive write-up.
Hope you enjoy, and let me know if you have any more technical questions I couldn't cover in the post!
Just a heads up: I'll be adding the 501's to the post soon.
I didn't even think about adding something on available sizes (which fit numbers come in short/long lengths or smaller/bigger waists). I should add that!
In the meantime, I checked to see which Levi's I've measured do come in length 36 for a size 32, and there are a handful of them:
* Levi's 505 Regular Fit
* Levi's 514 Straight Fit
* Levi's 527 Boot Cut Fit
* Levi's 541 Athletic Fit
Unfortunately there isn't too much variety, given that the 501's are also a regular seat/straight leg, but the 505/514's should fit a bit different and come in 32x36, so they might be worth a try!
They make comparisons to FlashAttention-2 when FlashAttention-4 has been out (even if they wanted to stick to Hopper class GPUs for whatever reason there's still FlashAttention-3). The two orders of magnitude claim look like they're for prefill not next-token decoding, which is a bit duplicitous. Long context extrapolation experiments typically go well beyond 2x context length. Etc etc etc.
I never said they should have a full public disclosure, but I do think sharing something of substance helps build trust and also get people excited.
Lastly, frontier labs have other incentives than to eek out every dollar and cent. Having the most capable models, not the most cost effective, is of significantly higher priority as OpenAI and Anthropic march towards IPOs. The same is not necessarily true for Google/DeepMind, and one can see from their public releases alone for some of their open weight models that this may be more of a priority for them today.