How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
Ironic that HuggingFace needed Chinese models to defend against it. But of course the spin of the leading firms will just be to point at their trusted access programs and demand that all dangerous activities, even if just defensive, happen via their APIs or be outlawed otherwise.
If that's the position they take then they really should be heavily regulated or nationalized. Cyberdefense against their own models dependent on their goodwill? Sure, but then they have to sell defense capabilities at subsidized rates with a limited margin. Would be very weird otherwise to take the world hostage with their models and then also sell the solution while demanding intrusive KYC.
Perhaps fortuitous timing for OpenAI that they can spin the fact that defenders have to resort to open Chinese models because OpenAI and Anthropic actively sabotage them with nerfed models into a nice message of making Huggingface part of the privileged group entitled to secure systems.
For me, harnesses are mostly sticky insofar as the model providers only allow you to use their subsidized plans through their own harnesses, unfortunately. But of course switching model + harness is an option.
While I'm sympathetic to the frictions of newly introduced AI and the fact that AI in healthcare, especially calls, can seem very uncaring, between the lines the article reads a bit like the typical union complaining about modern tech that reshapes their work, given the multiple mentions of protests, nurses union, etc.
Given how healthcare is one of these sectors that seems to relentlessly resist efficiency increases and is the prime example of Baumol's cost disease, I think any developed country with a costly healthcare system needs to do these AI experiments. The current versions will be shit, but the only way out is through if you still want to provide affordable care.
I honestly have no doubt that AI going forward will be able to do a good job at triaging via calls and also being empathetic about it. But of course it needs careful experimentation.
How do you couple them together efficiently? The nice thing about Codex or Claude is that the delegation or multi agent workflow capabilities are just built-in.
Do you link one with the other as a skill or mcp or so?
Interesting how the prevalent opinion until yesterday seems to have been that OpenAI & Anthropic are irreversibly ahead and now with xAI and Meta at least delivered something that's competitive with useful models and cheap too. Granted, the narrative that the two leading labs are ahead still holds with Fable (and perhaps an upcoming GPT6), but it's not as over as common knowledge by the opinion leaders would have us believe.
I just think that Anthropic's usage of the word "classifier", which implies a minimum level of intelligence, was very misleading. Fact is, you cannot use Fable for anything remotely connected to even elementary school biology or medical topics. There is no attempt whatsoever to distinguish between legitimate and dangerous tasks, except an extremely broad and non-specific rejection of anything related to security or biology.
Aren't they already very cognizant of handwringing like yours? Their article mentions various safeguards and actively steering the model away from being emotional companions and so on. It's a far cry from the OpenAI two years ago or whenever it was when they were entertaining the idea of allowing/enabling adult conversations with their models.
I personally think this is a moralistic regulatory overreach. And they definitely do that due to political pressure too, since there are various bills around the world in various legislatures that want to regulate AIs giving useful advice and being too personal to talk to.
So you can rest assured, I think, at least in that regard. The AI disempowerment will come to us anyway, just in a more sanitized corporate form.
Seems to be a very bad mechanism to ensure democratic control of the technology. There must be better ways, even naively assuming that OpenAI is somehow genuine about wanting to broadly share its stake in the future.
xAI has shown this to be quite lucrative. And it seems to even make some sense - if the contract can be ended on relatively short notice, you basically have the capacity on stand-by if you ever need it yourself (accumulating GPUs is not trivial), but can monetize it if you don't need it.
Though it's probably a bad sign generally that you can't capitalize on all the GPUs you've acquired.
That was my impression in the past as well, but by all accounts pangram works well, at least at the moment.
It also seems quite plausible that it can be made to work by training on a lot of model output. Most of us already have become very sensitive to the various idiosyncrasies of model writing, after all. They have a very distinctive style.
Yeah, I also got suspicious and checked it with Pangram. Sadly 100% AI. Perhaps it still has good points, but my heart drops whenever I sniff the AI prose. I can just query Claude or ChatGPT myself, you know?
The EU simply doesn't have a proper common market, especially when it comes to capital. Having more people than the US and a big economy in aggregate doesn't matter much if you can't efficiently pool resources. Could we in Europe have 100 billion fundraises for a new lab? If not, then it's over and you can give up.
Which is a mistake the EU makes again and again. If you put onerous requirements on everyone, this means that the most well-capitalized firms will be able to shoulder the regulatory overhead the easiest. But who can't? New European startups. This already killed part of the tech sector with the GDPR while Google and Meta just hire 100 lawyers and are done with it.
Wonder if the whole cyber paranoia leads to their models ultimately generating less secure code. After all, if it has the ability to generate safe code, it would imply that it knows something about cybersecurity, which could surely be used to hack all the banks in the world.
So it's like Claude Cowork for Science, i.e. for less tech-savvy users? I would imagine scientists with some coding background might just prefer to use Claude Code normally and integrate it with their stack of choice, but perhaps the comfort and ease of use of Claude Science still wins out.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.