> it seems that nontechnical POs have captured the software lifecycle
Yes, this dynamic has indeed been screwed up in many teams!
In my perspective, you can’t be a good product manager for a technical product without being somewhat technical yourself. You simply won’t understand your product well enough.
The only dynamic where that works is if the product manager acts more as a supporter and challenger to a team lead by an EM or technical lead, but then product manager is probably not the right title.
The big question is if they manage to keep on their enterprise consumers. Vowing changes nothing, they will be looking at the actual numbers.
Dissatisfaction is kind of expected (my cost goes up 2 orders of magnitude and I already cancelled since there are better options at market rate). Complaining without change won’t matter.
Read-only access to all non-sensitive code is how things should be. Huge engineering culture and productivity booster. It’s also very useful to keep each other honest (I’ve found so many “interesting” things hidden away in organizations with tight read access restrictions).
swe-bench is a standardized evaluation suite so that's why I'm asking - hopefully there are well-defined criteria on whether this is an open/closed book benchmark.
As I understand it, it is designed to evaluate the LM itself and not agentic systems with online access (very high likelihood of unintentional cheating/solution leaking). The paper and docs are not super clear on the concrete requirements (although reproducibility is emphasized which goes against online access). So I was hoping for someone with more familiarity to chip in.
Obviously not a problem for internal evaluations, but for fair scoreboard submissions it matters. It's not a matter of whether internet searches are useful, but rather what the benchmark is intended to benchmark.
I really really want to like local AI, but I highly doubt it will see wide adoption for a long time.
The additional up-front cost for hardware designed to run an LLM in addition to normal workload is unlikely to be accepted by most consumers.
The scale will be very constrained (like Apples on-device models which are small, heavily quantized, and have a small 4K token context window). It’s also terrible for battery life.
AI as it is implemented today is simply just computationally expensive and unless you put in dedicated hardware (like the ANE) for only this purpose - a large cost driver - I don’t really see it getting large scale adoption.
Companies will probably need a server-backed solution as fallback if they want reasonable user experience, so why even invest in diverse hardware support.
My thought exactly! First the usage limits + model limitations and now fundamental change to the billing. Hope some consumer watchdogs are looking into this!
I’ve worked most of my career in US tech satellite offices and I have not experienced EU team members to be less productive than US team members, nor spend less time on work (if anything, more really since they also need to be available for US time zone overlap).
It’s true there are chill jobs here, as there are in the US.
But ambitious people tend to work as much as ambitious US people (and it’s really more like 40 hours work weeks - 39,5 where I live since lunch is not work time). But again, many are not really counting, it’s just a full time job.
Vacations (typically 3 weeks summer holiday and additional weeks to distribute over the year) does create longer time on skeleton crew. Skilled tech labour is also cheaper so you can just hire more to make up for it.
Transparent screens doesn’t make much sense for consumer TVs (I know the article indeed points to other use-cases). You still need a black background to facilitate display of black content.
More likely crash looping of so many VMs overloading some system with insufficient back pressure, possibly combined with unfortunate cluster management scheduler behavior at this scale of crash looping (e.g. too eager to retry scheduling instances, maybe even on new hosts which causes more infrastructure load).
I believe Show HN is also what is suggested in guidelines for non-YC startups that are not allowed to use Launch HN (although some text along with the post would have been nice).
Exactly. It’s just leveling the playing field. I’ve been doing generated cover letters based on cv, job post and a few other data sources with manual review + adjustments with a very decent callback rate.
At least in danish, we have “bedstemor” which behaves likes “oldemor” in being the mother or fathers mom so it’s really that form that is consistent with the system and not the parent-specific forms.
> I find it very difficult to find meaning in a large portion of the jobs available today. Most workers are just another cog in the wheel. The system is so large that one cannot directly appreciate the effects of their work, or even know whether they are positive at all.
My experience is that the same job also has a large range for meaningfulness depending on how well leadership manages to facilitate it.
Yes, this dynamic has indeed been screwed up in many teams!
In my perspective, you can’t be a good product manager for a technical product without being somewhat technical yourself. You simply won’t understand your product well enough.
The only dynamic where that works is if the product manager acts more as a supporter and challenger to a team lead by an EM or technical lead, but then product manager is probably not the right title.