Any elite selection process has high risk for not selecting the best people.
Either due to corruption or, even more likely, institutional biases.
Malcolm Gladwell has a few episodes of revisionist history that looks at some US universities who fall into this trap.
Even with all that said; the chances of a selection process for 17/18 year olds being able to correctly select the best future adults in their field is very low.
Another good example; the number of young geniuses or prodigies (whether maths, chess, acting etc.) who don't make anything of themselves.
Congrats Hannah. I'll pile on with my own anecdote; I've run an internal conference at work for several years and we always book an external speaker.
She remains my (close) second favourite[1], with a choose your own adventure talk about algorithms. It was genuinely great, and warm, and thoughtful (a lot of it was about the risks of algorithms, which is very relevant for us).
This was... 2021 so mid COVID (yep it was remote) and so just before she became really big. We've worked with a lot of external speakers since, many of them very good or very slick, but she was the only one that was clearly both brilliant at her topic(s) and also an exceptional communicator.
[1] For the record my favourite is Clifford Agius who is some guy I saw at a local conference who in his spare time from being a commercial pilot builds fully robotic arms for kids - he even lets them give the finger. Great guy and refused to charge us more than a few hundred quid.
As I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO.
What sufficiently hard, but useful, problem would you ask the model for?
Agreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it.
The problem isn't the test, its that is a public test.
Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.
That's totally fair and things may change. For me its the history and the fact I can come back to it.
If I am honest I believe my final solution will be a combination of Open Claw, a custom knowledge wiki based on Wikmd. I just need a good all for Claw with history that is as good as gpt
Edit: and context too. It inferred my energy supplier from previously chats and so when I just asked a pertinent question it referenced their policy. Admittedly Google will have way more context if they get the product right.
For example; ChatGPT is replacing my Google searching. Not necessarily because it's better, or because it's summaries are better than Google (I find them subjectively better but it's not clear cut).
But because the app has a nice history; can ask a relatively complicated question and go do something else and then come back to it, ask a follow up. Etc.
None of that is specifically an AI benefit, but it's a workflow that really helps, well, flow.
Obviously I know nothing about your product, so completely uninformed!
But as an enterprise buyer $50/m and $10K/m is the same bucket in terms of cost. No one will blink until around 100K, depending on what it is.
(The point I am making is; as an enterprise buyer I absolutely know how annoying it is for me to turn up and go "this random regulation, we're interpreting it in this highly specific and unique way, and we want it asap". Hence willingness to pay down that inconvenience)
Your getting that interest because it looks like a steal. Ultimately those businesses couldn't care less about $50/m (except to chance it) but they want - or even need - the enterprise terms.
They will pay $50 for your product... And probably $950 for the terms.
(Not saying that would have been the right thing for you but my advice to folks who find themselves in this position is always 20x or 40x the price - if that is enough to make it worth your bother, then go for it. Good chance theyll pay)
I agree with you, but seperate point in many respects - the conversation was about replacing existing robust medical infrastructure.
I fully agree that AI could extend access; but to build on what others have said too, lack of physical diagnostics is an issue as is the lack of physical tech infrastructure.
Doctors make errors all the time though, so the real argument is about the error percentage. If AIs is lower then it's safer (but it's hard to have that convo, I recognise).
Besides; this article was about diagnosis not prescribing. It's pretty obvious, I think, that diagnosis is one area where AI will perform extremely well in the long run.
I think there are two metrics; the first is outright misdiagnosis, which studies put between 5 and 8% in US/Europe. That's a meaningful number to tackle.
Secondly; overdiagnosis. Where a Dr says on balance it could be X on a difficult to diagnose but dangerous problem (usually cancer). The impact of overdiagnosis is significant in terms of resources, mental health, cost etc.
What matters ultimately is the system achieves your goals. The clearer you can be about that the less the implementation detail actually matters.
For example; do you care if the UI has a purple theme or a blue one? Or if it's React or Vur. If you do that's part of your goals, if not it doesn't entirely matter if V1 is Blue and React, but V4 ends up Purple and Vue.
I meant that frame very deliberately. Use of the word AI is misleading people that LLMs are intelligent.
They model what looks like intelligence but with very hard limits. The two advantages they have over human brains are perfect recall and data storage. They are also faster.
But the brain is vastly more intelligent:
- It can learn concepts (e.g. language) with an order of magnitude less information
- It responds in parallel to multiple formats of stimuli (e.g. sight/sound)
- LLMs lack the ability to generalise
- The brain interprets and understands what it experienced
That's just the tip of the iceberg. Don't get me wrong: I use AI, it is by far some of the most impressive tech we have built so far, and it has potential to advance society significantly.
But it is definitely, vastly, less intelligent than us.
I just feel this is a great example of someone falling into the common trap of treating an LLM like a human.
They are vastly less intelligent than a human and logical leaps that make sense to you make no sense to Claude. It has no concept of aesthetics or of course any vision.
All that said; it got pretty close even with those impediments! (It got worse because the writer tried to force it to act more like a human would)
I think a better approach would be to write a tool to compare screenshots, identity misplaced items and output that as a text finding/failure state. claude will work much better because your dodging the bits that are too interpretive (that humans rock at and LLMs don't)
If a manager is handling (almost) all disputes of all sorts, then they will fundamentally lack authority to enforce an outcome on a real dispute. They simply are too involved because resolution requires you to take some sort of side.
If my children won't speak to each other I will refuse to be the go between because I become a proxy for one to the other. If one then punches the other they won't respect my perspective that this was wrong because I've set myself up as the proxy for the others feelings.
If you need a manger to resolve the above example, the org is broken and the engineers are poor engineers.
http://www.errant.me.uk/
http://twitter.com/errantx
http://www.errant.me.uk/blog/2009/10/pro-tip-tell-us-exactly-what-your-offering/
If you want to chat or comment or whatever I always welcome emails
[ my public key: https://keybase.io/errant; my proof: https://keybase.io/errant/sigs/m7EZLmQV9GiiOriXznJq8uH1p6RQ3t54rx1yRf4FdOs ]