Yeah, this is just slop. No benchmarks, no concrete case studies, just some vibecoded "platform" to finetune models on your own traces.
Which is an idea that has some value, but also some weaknesses. And this implementation of it isn't forthcoming with that concept. You have to really dig in to understand what they're even talking about.
ARC-AGI3 doesn't seem like a great benchmark to me in the first place. It assumes a lot of human like tendencies which an AI either shouldn't or wouldn't have. Particularly in the genre of "gameplay" where unspoken assumptions from prior games inform our understanding of rules.
You've got a good point, but they might spend a lot on the software side too, and any kind of weapons manufacturing in America will be expensive compared to other countries.
There aren't many local image editing models, so this is a welcome contribution! It's very compact considering the good demo results from their webpage.
My biggest gripe against eSIMs is that switching from one phone to another is difficult. I wish I could simply send a file to the other phone and get along with things.
Seems like this is particularly good at instruction following, but not as strong at coding as others. It's always great to get more diversity of open weight models though! I'll need to test this out to see what its "personality" is like.
> What is it that matters for "immediate mode"? Is it that the UI renders everything every frame? No. It's that you build the UI by describing what it should look like everyframe, based only (or mostly) on the data. This is why React won...
I don't think that's why React got so popular. React popularized unidirectional data flow, which is different than immediate mode rendering. This readme file seems to conflate the two of those.
Now that I think of it, couldn't one argue that React itself is a retained mode UI, since it choses which components to re-render and which not to?
I'm sorry, but I agree with the wiki editors in this case. Odin is obscure. The author of this blog post seems to think it's well known, but I don't think that's substantiated.
That understates how difficult it is to get to the level of performance they attained. The fact that it surpasses DeepSeek v4 in most ways shows that they accomplished some great work in this space.
I think you're missing the point. Everything you said is theoretically correct, but the parent comment was talking about the concrete circumstance of pentesting with the top models today.
Let's just take GPT 5.5 and Opus 4.8 as an example. Both are worse than Mythos 5, but they're capable of quite a bit when the guardrails are lifted and they're paired with a skilled human operator. They more than "good enough" to reach the same result with the addition of some human effort.
It is strange, huh? But the hype cycles around these models often ignore good contenders. Xiaomi's MiMo-V2.5 Pro was doing really well and didn't get much hype either.
Which is an idea that has some value, but also some weaknesses. And this implementation of it isn't forthcoming with that concept. You have to really dig in to understand what they're even talking about.