I was curious to see how open weight models would do on this task so I passed in a screenshot of your source of truth and here's what 2 of the best code-generation models that allow image inputs do:
Interesting to see Anthropic now downplaying the new vulnerabilities that Mythos discovered:
> We reviewed a demonstration of this specific technique being used to identify a small number of previously known, minor vulnerabilities. These vulnerabilities all appear relatively simple, and we have found that other publicly-available models are able to discover them as well without requiring a bypass
You could use Hugging Face's Inference API (which supports all of these API providers) directly making it easier to switch between them, e.g. look at the panel on the right on: https://huggingface.co/openai/whisper-large-v3
(b) growth is still the right metric imo. for example, lots of open-source libraries (which by definition are available to use for free) measure and optimize for growth
Lucky to be part of this company and seen this strategy work close up. One thing I'll add is this "decentralized" approach applies to all Hugging Face teams, not just the developer advocacy team. Just to give an example, there's no central comms team at Hugging Face, every team (usually the engineers who work on the product or features) does their own comms across the channels they think work best. That means there's lots of experimentation and most of our hires tend to be generalists who are comfortable wearing many hats.
Inkling (not too great): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Kimi 2.7 (really well, esp. note that this is the predecessor model, not the latest Kimi3): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Here's how I tested them: https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...