I work in Trust and Safety (not at Meta). We've never 'randomly' shut down an account and any action involving deactivation or deletion goes through thorough human review with no exceptions.
The lack of respect for the end-user is squarely a Meta problem, not an industry problem.
This is now the 5th comment saying the same thing, so I'll respond. I'm aware of these and they were terrible. In a just world, they would get as much if not more media attention.
The difference is the public nature of the execution. That is what makes it more similar to, say, Colombia or Venezuela _to me._ Within the context of 'magical realism', it is the perspective and mass dissemination of the violence that heightens that feeling.
Going back to the original topic, there is a reason that most of 100 Years of Solitude's pivotal moments happen around the staging of public executions (and not so much the off-screen violence, of which there is some but it's not focal).
On Sunday, I was talking a Mexican friend about how politicians get killed in our countries (Colombia, Venezuela, Mexico). Just in June, presidential hopeful Miguel Uribe was shot and killed in Bogota. In the head, in front of a crowd.
I remember being grateful about how that doesn't really happen in the US (Trump being the most recent, but he survived). I guess I was wrong... and, in that case, Garcia Marquez might agree with you.
The most direct, non-marketing, non-aesthetic summary is that this model trades off a few points on 'fundamental benchmarks' (GPQA, MATH/AIME, MMLU) in exchange for being a 'more steerable' (less refusals) scaffold for downstream tuning.
Within that framing, I think it's easier to see where and how the model fits into the larger ecosystem. But, of course, the best benchmark will always be just using the model.
I used to work at a drug discovery startup. A simple model generating directly from latent space 'discovered' some novel interactions that none of our medicinal chemists noticed e.g. it started biasing for a distribution of molecules that was totally unexpected for us.
Our chemists were split: some argued it was an artifact, others dug deep and provided some reasoning as to why the generations were sound. Keep in mind, that was a non-reasoning, very early stage model with simple feedback mechanisms for structure and molecular properties.
In the wet lab, the model turned out to be right. That was five years ago. My point is, the same moment that arrived for our chemists will be arriving soon for theoreticians.
LLMs are really annoying to use for moderation and Trust and Safety. You either depend on super rate-limited 'no-moderation' endpoints (often running older, slower models at a higher price) or have to tune bespoke un-aligned models.
For your use case, you should probably fine tune the model to reduce the rejection rate.
I feel like the bash only SWE Bench Verified (a.k.a model + mini-swe-agent) is the closest thing to measuring the inherent ability of the model vs. the scaffolding.
Papers have been doing rollouts that involve a model proposing N solutions and then self-reviewing to choose the best one (prior to the verifier). So far, I think that's been counted as one pass.
Are these simulations shared between your customers, or are you building bespoke environments per client/user? How does the creation of environments scale?
https://www.nature.com/articles/s41598-018-38096-z