Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.
in the short term, ai has been job-creating and unemployment remains quite low. The technology has mostly been augmentative (like every other technological revolution beforehand)
if (big if) the technology continues to improve at the current rapid pace, then we would need to organize welfare for the bottom % of population that cannot compete with ai or robotics in any feasible way, but that kind of presupposes insane productivity growth such that living off 'welfare' is a life of luxury compared to the average life right now
This is one of the best writeups I've seen of this
a lot of the model's constraints come down to how they are RLed. Discussions online would be a lot better if everyone understood how the labs train the models in a high level (or did a lil data labeling)
I'd wager you get similar results if you gave people a version of Google search that purposely gave you bad results. Like, it's framed as an assistant / lookup tool - is it so surprising that people tend to trust it more? Especially since the participants are likely used to using full-powered models and the researchers give them a purposely gimped one (lol)
People are acting rationally when given AI tools to lookup information, their first consumer use case was as a super-powered Google Search
You are correct that LLMs are trained on existing proofs but hiring researchers to solve unsolved problems is just unrealistic, both in terms of how none of the mathematicians simply came out and took credit for their own discovery or exposed this, and how training sets are not easily memorized (rather, the meta techniques are learned).
OpenAI just has better training methods and techniques for pure math over Anthropic, it’s one of their biggest strengths
The top % of YC founders are legitimately quite talented, and high agency/entrepreneurial talent goes a long way in these companies, even if its not in deep infra or research
i dont think the comparison to traditional vc saas is very good
yes, in traditional venture you want cost per marginal user to decrease and leverage your platform at scale
but improving llms shifts the frontier of their capability and unlocks entirely new use cases. so far, every mega training run has resulted in a model that has paid itself off profitably fully loaded. perhaps the TAM of intelligence has no ceiling?
not to say that we shouldnt be investing in efficient models, but the efficiency comes after we create another mega shoggoth that we can make more efficient
Id say the main difference is FDEs post-engagement need to drive product strategy back into the platform (non trivial ask)
you typically see FDE-driven companies' products be 'assembly' driven and very deep into integration, as they figure out the optimal primitives that assemble into the shapes required to solve new customer problems
your FDEs shape your product strategy, and should be considered R&D. after making sure a customer deployment is successful (by any means necessary btw, even if it means building new systems outside of the product), the crucial next step is to drive the product improvement with PMs and core software engineers after contact with reality. this was a pretty radical idea from palantir in the era of saas
if you only do step 1 you're basically just solutions engineers / mckinsey, and if you only do step 2 with no customer learning to your product you don't improve your platform for all the other customers. the pain becomes the moat
There's a reason why this echelon of companies comp FDEs much, much more than services businesses is because you're trying to find engineering + product + customer facing in one (knew people making 200k+ 5 years ago as new grad FDEs, and the same flavour at the labs is 500k+ easy)
that being said the role has evolved a lot over the years, and depending on the company it could be indistinguishable from solutions eng, or sales eng, or even dev rel.
sadly with all the labs benchmaxxing I feel like you just have to try the model for a while to really evaluate how good it is, especially for each individual use case
100%, if someone from a no name school does well on the interview I’ll happily recommend them to be hired. However idk how HR filters resumes, and they likely use certain heuristics to try and minimize false positives
Do you need to do tier 1 and 2 work before tier 3?
If they are structurally different, and there’s a way to train people directly into tier 3, then it doesn’t seem unreasonable to automate t1 and t2 as from my experience the vast majority of the tickets are either simple or repeated workflows. Taking the idea to the limit, you’d automate all tiers, and have the ai escalate to the individual teams within the company for any truly meaningful edge cases
I feel sort of the same about SWE, which is much more complex, but juniors can ostensibly grow into seniors with AI
model routing in this case is cross-provider
Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.