I really like how you put it down! This is how I'd have imaged to be the Apple car, not a Ferrari. The "irony" of people defending this monstrosity are saying thing like "this is the type of comments of someone who cannot afford this", but they are completely missing the point. If it doesn't making people dream about it, then it's not worth of being a real Ferrari product.
> Don't focus on what you prefer: it does not matter. Focus on what tool the LLM requires to do its work in the best way.
I noticed that LLMs will tend to work by default with CLIs even if there's a connected MCP, likely because a) there's an overexposure of CLIs in training data b) because they are better composable and inspectable by design so a better choice in their tool selection.
HumanSignal | https://humansignal.com/ | REMOTE North America, South America, Europe | Full-time | Engineering roles
We created Label Studio (https://github.com/HumanSignal/label-studio/), which has quickly become the most popular open source data labeling and AI evaluation platform with 350K+ users around the world and millions of annotations each month, alongside a community of thousands of data scientists and ML engineers sharing knowledge and working to advance AI.
We're a remote team full of people passionated about open source and AI. We are very pragmatic and strong team players.
We are looking for multiples roles to support the growth of Label Studio:
I think they are exposing how fragile and vulnerable in reality they are, and I wonder when it will happen that a group of highly motivated individuals will organize to create a truly community driven distilled models.
I think they are exposing how fragile and vulnerable in reality they are, and I wonder when it will happen that a group of highly motivated individuals will organize to create a truly community driven distilled models.
> We find testing and evals to be the hardest problem here. This is not entirely surprising, but the agentic nature makes it even harder. Unlike prompts, you cannot just do the evals in some external system because there’s too much you need to feed into it. This means you want to do evals based on observability data or instrumenting your actual test runs. So far none of the solutions we have tried have convinced us that they found the right approach here.
I'm curious about the solutions the op has tried so far here.
> And I’m saying this as a Swede. Buy German cars, specifically within the Volkswagen auto group (Audi, VW, Skoda etc) if you want reliable quality.
I own a 2020 BMW with an electronic gearbox, which broke at around 80k km just a couple of months after the warranty expired (yeah I know!). It was a bit of a headache going back and forth with BMW to request a free repair. Fortunately, the headquarters agreed to cover the cost, and they installed a refurbished electronic gearbox. I was quite relieved that I didn’t have to pay about €10K out of pocket!
All that to say that I wouldn’t call BMW particularly reliable in terms of quality these days, but their customer support was decent, at least in my case.
Ultimately those are tools and I think the goal is to educate students to use them properly. Also because I don't expect the knowledge paradox to disappear anytime soon with these models.
I'd have agreed with you, if the principles would be different. But what was showed in the content is EXACTLY what those tools are doing today. Actually those tools are way more powerful and considering & covering way more scenarios.
> There’s nothing wrong with starting from scratch or rebuilding an existing tool from the ground up. There’s no reason to blindly build from the status quo.
Generally speaking all the options are ok, but not if you want to have something up as fast as you can or if your team is piloting something. I think the time you spend to vibe code it is greater than to setting any of those tools up.
And BTW, you shouldn't vibe code something that flows proprietary data. At least you would work with co-pilots
> Q: What makes a good custom interface for reviewing LLM outputs?
Great interfaces make human review fast, clear, and motivating. We recommend building your own annotation tool customized to your domain ...
Ah! This is a horrible advice. Why should you recommend reinventing the wheel where there is already great open source software available? Just use https://github.com/HumanSignal/label-studio/ or any other type of open source annotation software you want to get started. These tools cover already pretty much all the possible use-cases, and if they aren't you can just build on top of them instead of building it from zero.
Recently started using Cursor for adding a new feature on a small codebase for work, after a couple of years where I didn't code. It took me a couple of tries to figure out how to work with the tool effectively, but it worked great! I'm now learning how to use it with TaskMaster, it's such a different way to do and play with software. Oh, one important note: I went with Cursor also because of the pricing, that's despite confusing in term of fast vs slow requests, it smells less consumption base.
Just finished FF VII Rebirth, which I'm considering exactly what a FF should be, with the exclusion of the last chapter's narrative that I didn't like. That said, next one is Clair Obscur, very looking forward to play it!
I don't have faith that this is something we can fix in the short term because most of us have been educated in a very competitive environment where individuals come first. I'm not saying that the opposite is good either, but we should find a balance in between. I also feel like that we are all becoming more disconnected, alone, and where the center of gravity is only ourself. Despite my premise, I still have some hopes for future generations, but unfortunately I think that things will get way worse before correcting.
> B) Demographics are now working against us instead of for us -- turns out everyone has decided not to have kids, which means an end to population growth, consumption growth, ergo hiring growth
Even in the case of population growth things will not look better especially because we are in the middle of a tragedy of the commons, we are exhausting the resources of the planet faster than what it takes to regenerate them. What's the plan for more consumption when there's nothing to produce? or that cost so much that only a few could afford that?
I guess that we will see more requests for data labelers that know coding from LLM providers to answer Stack Overflow like of questions in order to keep their model up to date.
> Sales is everything in B2B software and always has been. Product-led growth in B2B has always been fantasy erotic-fiction outside of chat/notes apps.
PLG creates the distribution to actually implement Sales effectively and at scale for B2B. Even those companies that originally were purely Sales-led now have a strong PLG component. PLG is far from be a fantasy, but now it's almost a must have motion to build distribution and long term healthy business viability considering also that Enterprise software is living a consumerization moment. Not to mention that the next generations of users and buyers buy and expect software to be different from the past, this is already happening.
HumanSignal | https://humansignal.com/ | REMOTE North America, South America, Europe | Full-time | Engineering roles
We created Label Studio, which has quickly become the most popular open source data labeling platform with a 250K+ users around the world and millions of labeled samples each month, alongside a community of thousands of data scientists sharing knowledge and working to advance data-centric AI.
We're a remote team full of people passionated about open source and AI. We are very pragmatic and strong team players.
We are looking for multiples engineering roles to support the growth of Label Studio:
I think that's a false myth, there are now EOR like Deel and Remote.com that make this super easy today.
There's also the "false contractors" route, i.e. hiring remote contractors but treating them like employees. Which is also a quite common setup among early
stage startups with a fully distributed DNA.