AIUI: my intuition is Alex and Alice are points in the distribution. They don’t think about their experience in terms of population statistics. They see their individual latency times, and use that as their sample. If t is low in their experience, great the distribution is low.
But for any t that goes high that they observe (which tends to be the case in a skewed distribution such as service latencies), it drags their impression of the distribution up, dominating the shape of that impression.
If we can just simulate the business accurately enough we can solve having to interact with the market… which is also trying to solve interacting with us… We just need to do it more accurately…
Something tells me people will still be in the mix here.
A cursory glance does make it appear like either a prolific individual, or a bot. The fact that the novel bears little relation to the analytics posts, which seem to bear the style of LLM prose, makes the whole thing fishy. Ironic given the subject matter of TFA
Revenue recognition for private companies generally is less precise than for public companies, which in the US are obligated to report under GAAP, which uses a different indicator than annualised revenue so public companies are comparable.
It makes sense to scrutinise Anthropic’s revenue in the lead up to IPO on those grounds; their AR figure simply isn’t comparable to revenue numbers from other firms.
However it doesn’t make sense to be sensational about this - iirc reporting GAAP revenue is a necessary condition of going public so the chickens will come home to roost one way or another.
Yep agree it looks like it’s taking the existing generated artefact, parameterising it within an inch of its life, exposing a pseudo WYSIWYG for the parameters and calling it a day with a few export options. Not a huge leap from what they’ve got already but it’s a clever adjacent step for sure. Same product new chrome.
You’re illustrating one of the points of TFA - a team that is equipped with the right tools to measure feature usage (or reliably correlate it to overall userbase growth, or retention) and hold that against sane guardrail metrics (product and technical) is going to outperform the team that relies on a wizardly individual PM or analyst over the long term making promises over the wall to engineering.
Without commenting about the frequency of negligence myself, I suspect at least that you and GP are in agreement.
I doubt GP is suggesting ‘go ahead and be negligent to feedback and guardrails that let you course correct early.’
Plugging the Cynefin framework as a useful technique for practitioners here. It doesn’t have to be hard to choose whether or not rigorous planning is appropriate for the task at hand, versus probe-test-backtrack with tight iteration loops.
Chiming in as Australian with no context on European situation. AFAICT the key drivers of cost inflation are to do with reconfiguring the electric grid to transfer power efficiently and reliably from plants that produce renewable energy. However, the grid is set up to do so from non-renewable sources. And you want to do it while smoothly operating the network. This is extremely hard. Doing so quickly therefore elevates prices. That’s the rationale I could imagine being the case in EU markets.
That’s possible, I suppose. I think @btbuildem was expressing a personal distaste for other uses of power, and an avulsion to the technology because of that. For example: labor camps.
Iiuc it wasn’t a comment about what the perfect lifespan is. It’s expressing a concern about how people in power might apply life extending technologies, like they do many other technologies, to exercise and entrench that power.
Or put differently: it’s a request, given limited resources let’s expend effort on a fairer society, not one with longer lived people.
Probably worth clarifying with GP what responsibility and accountability they’re referring to.
Where I live, if an engineer signs off on a bridge design and the bridge personally collapses, they are personally liable for harm done to folks on the bridge. As far as I’m aware software engineering does not have something like that.
For one - I’d say scoped API tokens that prevent messing with resources across logical domains (eg prod vs nonprod, distinct github repos, etc) is best practice in general. Blowing up a resource with a broadly scoped token isn’t a failure mode unique to LLMs.
edit: I don’t have personal experience around spending limits but I vaguely recall them being useful for folks who want to set up AWS resources and swing for the fences, in startups without thinking too deeply about the infra. Again this isn’t a failure mode unique to LLMs although I can appreciate it not mapping perfectly to your scenario above
edit #2: fwict the LLM specific context of your scenario above is: providing examples, setting up API access somehow (eg maybe invoking a CLI?). The rest to me seems like good old software engineering
> That’s exactly what it means to hit a wall, and exactly the particular set of obstacles I described in my most notorious (and prescient) paper, in 2022. Real progress on some dimensions, but stuck in place on others.
The author includes their personal experience — recommend reading to the end.
But for any t that goes high that they observe (which tends to be the case in a skewed distribution such as service latencies), it drags their impression of the distribution up, dominating the shape of that impression.