This is a very interesting point, in some ways the implicit belief is that we just need to get beyond the 700g limitation in terms of scaling LLM models and we would get human intelligence/superintelligence.
I admit I didn't really get the body/brain analogy, I would have been better satisfied with a simpler graph of brain weight to intelligence with a scaling barrier of 700g.
So basically preventing dead latents from occurring and whenever they do occur to possibly reviving them through the use of auxiliary loss term in the loss function? Thanks btw
Coming from a layman's perspective, a genuine question regarding:
"Implements SAE training with auxiliary loss to prevent and revive dead latents, and gradient projection to stabilize training dynamics".
I struggle to understand this phrase "to prevent and revive ", perhaps this is simple speak to those that understand the subject of SAEs, but it feels a bit self contradictory to me, could anyone elaborate?
What do you think about having codebuff write a parser for javascript? Something that is specifically built to enhance itself that goes beyond the regular parsers and creates a more useful structure of the codebase to be then used for RAG for code writing?
This would be double useful as a great demo for your product as well as enhancing your product intrinsically. For example the new parser can not only build the syntax tree but also provide relevant commentary for each method to describe what it does to better pick code context.
Based on what I've read it is allowed for a non profit to own a for profit asset.
So I'm assuming the game plan here is to adjust the charter of the non profit to basically say we are going to still keep doing "Open AI" (we all know what that means), but through the proceeds it gets by selling chunks of this for-profit entity, so the essence could be the non-profit parent isn't fulfilling its mission by controlling what openai does but how it puts the money to use it gets from openai.
And in this process, Sam gets a chunk (as a payment for growing the assets of the non-profit, like a salary/bonus) and the rest as well....?
If the system took 3 days to solve a problem, how different is this approach than a bruteforce attempt at the problem with educated guesses? Thats not reasoning in my mind.
There is a difference between Peak Oil "Output" vs Peak Oil "Demand".
The article is about demand and not production capacity which won't matter once the demand starts to go downhill... will never be zero but will not be something that can prop up regimes that have sucked the life out of their people for more than half a century.
Ah I see you like to generalize. This plane was lucky that it didn't burst into flames as soon as everyone exited or worse yet while they were still exiting.
There are multiple seats per row and only the first person of that row may benefit from the concurrency, as each person gets into the isle for that row they will add those seconds and that is sequential not concurrent.
Also you likely haven't been in a trampling crowd situation, bag on a chest doesn't occupy zero space.