Label studio is fine if it covers your need, but in many cases the core opportunity in an eval interface is fitting in with the SME’s workflow or current tech stack.
If label studio looks like what they can use, it’s fine. If not, a day of vibecoding is worth the effort to make your partners with special knowledge comfortable.
I’d you’re interested in using one of the LLM-applications I have in prod, check out https://hex.tech/product/magic-ai/ It has a free limit every month to give it a try and see how you like it. If you have feedback after using it, we’re always very interested to hear from users.
As far as fine-tuning in particular, our consensus is that there are easier options first. I personally have fine-tuned gpt models since 2022; here’s a silly post I wrote about it on gpt 2: https://wandb.ai/wandb/fc-bot/reports/Accelerating-ML-Conten...
I’d you’re interested in using one of the LLM-applications I have in prod, check out https://hex.tech/product/magic-ai/ It has a free limit every month to give it a try and see how you like it. If you have feedback after using it, we’re always very interested to hear from users.
We allude to 2 when talking about using explanations first, but I totally agree. One minor comment is explanations after can sometimes be useful for understanding how the model came to a particular generation during post-hoc evals.
Point 1 is also a good callout. I added something on this for llm judge but it’s relevant more broadly.
I’m really into this gps game: https://wandrer.earth (I’m not the creator, just a fan)
It encourages you to walk/run/bike all the roads in a town/city/county and it tracks the percentage.
It’s a fun and creative way to go see parts of your area you wouldn’t otherwise. I’ve used it to walk 97% of the miles in Berkeley and seen so many funky little areas. I also really enjoy how much it has increased my walking over the last few years.
The database is just filled with garbage that is unusable
for this project: Fonts that are completely illegible, fonts
that are missing most of their characters, fonts with millions of control points, Comic Sans MS, fonts where every
glyph is a drawing of a train, fonts where everything is fine
except that just the lowercase r has a width of MAX INT,
and so on.
The themed time-intervals suggestion is not unreasonable, but this article is completely devoid of any evidence to suggest its success besides... checks notes Elon Musk running several companies.
This is sorta true, but not quite. Leland McInnes and John Healy (the creators of UMAP), do in-fact have an amazing paper on HDBSCAN, but it's not inventing it. In their paper, https://arxiv.org/pdf/1705.07321.pdf, they introduce AHDBSCAN which is a great extension of HDBSCAN to dramatically improve it's performance.
Their work is great but just wanted to save people a google in case they were interested.