It seems clear that LLMs significantly change threat model math, but this observation alone does not explain how or why; the asymmetry that you’re describing is a property of pre-LLM software as well.
Each of those examples varies widely, and I don't think most people would treat each of them the same way.
In general when the stakes are higher and the ambiguity of outcome is less clear, secondary signals become more important.
Concretely: I don't give a shit if my housecleaner doesn't make their own bed as long as they make mine; the outcome I need is easy to verify and the stakes are fairly low so the secondary signal doesn't matter very much. Conversely, I care a lot if the therapist I'm relying on to help me manage my depression is visibly unable to manage their own; the outcome I need has a slow feedback loop and the stakes are high so I'm much more likely to rely on secondary signals like "is this person able to manage their own mood successfully?"
> It doesn't know if 1 person made all those requests, or N.
FWIW this is highly unlikely to be true.
It's true that the upstream provider won't know it's _you_ per se, but most LLM providers strongly encourage proxies like OpenRouter to distinguish between downstream clients for security and performance reasons.
I took the trouble of logging in to HN to call bullshit on this whole thread. The reason these teenagers aren’t building standalone apps is generally because they know that web apps are the future. Moreover, there are dozens of new travel planning startups every year (YC regularly funds them!) and they seem to discover the same thing that Microsoft did—the real reason they stopped updating this magical app—which is that everybody wants better travel apps but nobody wants to pay for it.
I’m sure someone will crack the code someday and I’m glad they will keep trying, but I refuse to accept the premise that $100 Apple developer accounts are their primary impediment.
For other public BigQuery datasets I’d certainly consider just adding them to the demo for free—which specific Bitcoin/finance datasets did you have in mind?
For private datasets, we’re looking at adding that functionality to the core Tabby service (that’s the SaaS this relates to). Please email for info!
Are you looking to set it up for another public dataset or your own private one? Either way, shoot me an email (on the demo page and also in my HN profile) and I’ll see what I can do.
Thanks for noting this! The functionality is definitely not perfect, and the opaque nature of the underlying model does not give much opportunity for tweaking. I suspect slightly altering your input might help drive better output.
The weather dataset has a bit of an unorthodox schema IMO, which gives the model more trouble than usual. That’s kind of the point of this demo, though: to what extent is generated SQL like this useful—despite its flaws—in the context of real-world datasets? Jury is still out :)
Nice! Yeah I have no doubt that a specialized model could beat this general one, although I find the output from the general one to be uncanny at times. Would love to hear your expert opinion on how they compare!
Yeah many rough edges indeed. The generated SQL is the plain output from GPT-3; I have not done anything to customize the model or validate syntax outside it, so the roughness is expected. No idea if folks will find value in this despite that, hence the demo.
In terms of actually tracking and charting this data, it’s pretty useful to have it all in one place and queryable by SQL—this is not usually the case by default!
This is a great overview of what to measure, at least for business metrics.
(Shameless plug…) For folks wondering how to measure, and also how to track product metrics for their SaaS, I wrote a thing: https://www.tabbydata.com/book