TLDR: the experiment asks for in-distribution responses and gets those.
The right answer here is to ask a LLM to create a scene similar in quality to those, but completely out of distribution.
I asked GPT 5.6 Sol to give me a pelican playing football on San Siro while smoking a cigarette, in AC Milan's t-shirt. While this sounds like higher complexity of a problem, the generations from current models often include additional details like scene composition, scarf, etc., I don't ask for, so I wanted to see what here is memorization vs. composition skill.
"write svg code of a fish playing football on san siro in ac milan's t shirt, with raybans on and a cigarette."
Try that on GPT 5.6 Sol, Fable, or whatever other model. It's chaos.
I am not expecting it, it's the Qwen team is claimingthey can do much harder tasks than this, like rendering a consistent page of a maths paper, or creating true to fact explainers.
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
A lot (all?) VCs charge some form of fees (typically capped at 20% of the entire fund, split in various percentages through 4 years investing, 4 divesting period). These fees often are only paid out only based on the actively deployed capital, and are not the only incentive: the main incentive is shares in gains (carry).
The reason they're based on actively deployed capital isn't that the LPs (people who give VCs money to invest) want them to deploy the money in a stupid way, but they definitely don't want VCs to get the fees if the money wasn't invested. Therefore, VCs:
1. Want to raise as much money as possible
2. Want to deploy as much money as possible
Ideally, as quickly as possible.
There's nothing fraudulent about the idea of calculating VCs fees in various scenarios.
There's however the extremely dodgy part of the portfolio companies paying their investor (VC) fees for anything. This is an obvious conflict of interests, and should never happen, but I personally know of multiple VC funds here in Europe (will skip the names to not get sued, lol) who base their entire operational model on funding shitty companies that have 0 chance of success, charging them for the office space and often "shared services" they provide. Unsure if this is a regulatory overlooking, or something that's deliberately legal, but IMHO shouldn't be. Probably they talked their LPs into agreeing to this on paper.
There have been numerous cases of sanctioning and wealth confisactions (Afrghanistan, Venezuela, Iraq, Iran, Libya), and "_now_ carries confiscation risk" is just factually incorrect. It has always carried such risk, this risk has materialized numerous times, and most importantly, no diversification is happening - literally nothing changed: https://data.imf.org/en/news/4225global%20fx%20reserves%20de...
By the way, good luck with trusting China, Russia, or other places to store wealth more than Europe. It's total ignorance to believe there's somewhere safer to store your money than the West.
"Pozsar’s argument: the moment Western nations froze Russian foreign exchange reserves, the assumed risk-free nature of these dollar holdings changed fundamentally. What had been viewed as having negligible credit risk suddenly carried confiscation risk."
This has nothing to do with dollar. Almost all of the confiscated currency was in Europe, and it has to do with invading your bank's ally. No country in the world ever assumed that's risk free, it wasn't being priced back then and isn't now, because nobody except from Russia is stupid enough to do that.
If you're adding some computational/problem breakdown/heuristic steps on top/instead of mathematical concepts, then you're doing the opposite of what the author proposes.
Scientific conensus in math is Occam's Razor, or the principle of parsimony. In algebra, topology, logic and many other domains, this means that rather than having many computational steps (or a "simple mental model") to arrive to an answer, you introduce a concept that captures a class of problems and use that. Very beneficial for dealing with purely mathematical problems, absolute distaster for quick problem solving IMO.
"Think in math, write in code" is the possibly worst programming paradigm for most tasks. Math notations, conventions and concepts usually operate under the principles of minimum description lenght. Good programming actively fights that in favor of extensibility, readability, and generally caters to human nature, not maximum density of notation.
If you want to put this to test, try formulating a React component with autocomplete as a "math problem". Good luck.
(I studied maths, if anyone is questioning where my beliefs come from, that's because I actually used to think in maths while programming for a long time.)
Agreed, but that's not in the C dimension of a first-layer embedding of a single token though, it's across the whole model and that's what I said in the comment above.
It's... really not what I meant. This requirement does not have to be relaxed, it doesn't exist at all.
Semantic similarity in embedding space is a convenient accident, not a design constraint. The model's real "understanding" emerges from the full forward pass, not the embedding geometry.
Language models don't "pack concepts" into the C dimension of one layer (I guess that's where the 12k number came from), neither do they have to be orthogonal to be viewed as distinct or separate. LLMs generally aren't trained to make distinct concepts far apart in the vector space either. The whole point of dense representations, is that there's no clear separation between which concept lives where. People train sparse autoencoders to work out which neurons fire based on the topics involved. Neuronpedia demonstrates it very nicely: https://www.neuronpedia.org/.
This is finetuned to the benchmarks and nowhere close to O1-Preview in any other tasks. Not worth looking into unless you specifically want to solve these problems - however, still impressive.
I mean again, even within a single sport, there are role differences, but the degree of fitness that you have to have to build 130kg muscle mass that JJ Watt has, and to run the 38.5 that Mbappe does... is just not the same level of fitness.
On the reverse, Mbappe has less strenght, and JJ Watt moves like a tank with 27km/h speed. If you compare them to elite strenght sports, they're both weak. If you compare them to sprinters, Mbappe is an amateur sprinter, and JJ Watt is disabled.
Sportsmen specialize in what they do, but NFL simply doesn't require versatility so they are good in fewer categories, and not very good in any. Soccer players are good in multiple categories, and by the virtue of being the more popular and competitive sport and having insanely bigger selection, occassionally very good in one or two (e.g. Bale, Mbappe, and other freaks of nature who are essentially sprinters).
Also, soccer player does not run 1/20th of marathoner runs. Elite wingers run just under _a third of marathon_ each game, of which 3km can be sprint.
Ugh, NFL players need to run on average about half of what a third tier soccer players do. At the extreme, a goalkeeper runs about the same as most running NFL player. For sprinting, the peak speed is similar, but soccer players run 2-3km of sprints during the game (1-1.5km for NFL). I'm not bringing NBA into this because it's just not a running sport altogether.
For explosiveness, top speeds of NFL and top 5 leagues in Europe are comparable, but more consistent for soccer players. They of course have to run with a ball next to their legs, rather than in hand, which makes it technically harder. For jumping, tall soccer players are closer to NBA players than to NFL (Tomori, Ronaldo, Lewandowski, etc, jump around 80cm).
In terms of agility, NFL and top 5 leagues is similar, about 3 seconds to 30km/h, but of course the best performing players in soccer are better.
So, with some similar parameters, soccer players do what NFL players do, but 3 times as long. That's the difference between "I can do this with a bit of a belly" and "I need to look like a god to even survive this game without getting a heart attack".
Edit: my point above wasn't that it's not physically difficult altogether, it was that these are not _elite_ sports in terms of physical requirement. Swimming, climbing, sprinting, soccer (mainly by the virtue of how professionalized it is), bicycle racing, that's physically super difficult. Basketball is super technical and relatively chill in physical requirements compared to these sports, and NFL is generally challenging but not nearly as much as the "top" sports, unless you specifically cherry-pick comparison to favor heavy, fast people. I chose rather versatile metrics that focus on input, e.g. how much you need to train to become fit enough.
Oh yeah, there are 4 slightly chubby guys in "elite" sports out of 10 000, therefore we got athleticism wrong.
Also, sorry to say it so directly, but none of these guys would have a remote possibility to (athletically speaking - not talking ability) play in third tier soccer in Italy or Spain, or any actually physically difficult sport (e.g. climbing). It's a rule in these sports that people look at minimum super fit, at maximum godlike, and the few exceptions that exist show extremely visible downsides (and it's clear they'd be better off being athletic).
For context, a picture of a soccer player considered unfit (constantly rated at something around ~70/100 physically in various rankings/video games/etc)
The right answer here is to ask a LLM to create a scene similar in quality to those, but completely out of distribution.
I asked GPT 5.6 Sol to give me a pelican playing football on San Siro while smoking a cigarette, in AC Milan's t-shirt. While this sounds like higher complexity of a problem, the generations from current models often include additional details like scene composition, scarf, etc., I don't ask for, so I wanted to see what here is memorization vs. composition skill.
"write svg code of a fish playing football on san siro in ac milan's t shirt, with raybans on and a cigarette."
Try that on GPT 5.6 Sol, Fable, or whatever other model. It's chaos.