This reminds me of an old problem I used to run into with duplicate data.
Duplicates are not always what they seem - duplicates have a relative quantity of how duplicate they are, for example - 100% duplicate identical is what it sounds like, but then what is a 50% duplicate? Well.. that could be a duplicate where semantic meaning that only a human is aware of (missing information) could be used to determine that these two datum are in fact duplicates in different forms, where 50% is identical, but the other 50% is semantically identical. Or 50% truly identical could, in some case, be "good enough" to be considered identical. These are just examples.
Why this example is relevant is because, when deduplicating data, we often have to consider which value(s) we keep as the true singular representation.
The problem is, long after the data is created, this may be impossible to do.
The solution? Often times it is no better than flipping a coin and guessing, i.e. randomly keeping one datum and discarding what we consider to be the duplicate. Or you can try to be fancy and merge the data or choose, based on some rule, what you assume is the data with the highest valid signal and discard or ignore the noise.
This happens every day in practical information systems with messy data.
I always come to the conclusion that in the end, one cannot truly replace missing signal precisely - one can only infer the correct signal. If the information was not recorded and or compressed or corrupted beyond recoverability, there is no way to get it back - only to infer or guess with some degree of confidence, and accept or reject the outcome.
This to me actually clarifies something deeply about the conversion between the real world to the digital one as data, and the conversion of data into information- it is a real physical process.
In my mind it makes clear that the recording of data (or generation of data) is a real physical process of transfer, from an instrument making observations to physical bits in a physical system, which become the raw data.
Without the raw, precise, complete data, your observation is lossy, and therefore everything above it is as well- e.g. information/interpretation, etc.
This isn't a negative. If stated correctly, the above just helps to explain how we imperfectly navigate through the real world with our tools.
All the same, we can do some pretty amazing things in the absence of data, even if some of it is "false" or "wrong".
I think that we humans still have a really tough time differentiating and classifying between public and private goods (and everything in between) from an economic sense.
Unfortunately, capitalism of today appears to find public goods and strong, public, healthy public goods to be undesirable investments. Yet, it seems inevitable that some things trend that direction.
I feel if we could get a better handle on what goods fall into which category, and find ways to make those things sustainable and even grow, the situation could get better.
I'm speaking purely from observations here but it also looks like in some ways, the "invisible hand of the economy" (i.e. Adam Smith) might sort these things out by itself, albeit, painfully.
We could look at the fact that LLMs were mostly created from the public domain (except when they were not), and the resulting products/services (expensive to create/operate, benefits a huge swath of humanity, has all kinds of externality issues) as something that perhaps should again almost be considered as a public good.
If we look at the struggles LLM/AI companies have had with legal structure choices, such as nonprofit, for profit, etc and pricing (e.g. do we make it accessible at a loss? Do we extract huge profits?) as a kind of moral dilemma of classifying this new economic good (LLM-based AI).
In the US, the government has been taking investment interests in some tech companies, and recent proposals even include establishing sovereign wealth funds around AI/LLM. I think this is further evidence that the system is trying to figure out how to classify these new economic goods and corporations.
So like.. conceptually kind of like memory mapped files on fast flash persistent storage, IIUC? Or maybe it's more like GPU-managed demand paging, caching and DMA? That could get you the capacity and better I/O characteristics.
I'm curious about how the unifying architecture is going to evolve between CPU/GPU having direct access to a singular pool of memory/storage also.
I also keep wondering when memristor technology might enter the ring, because as I understand it, it would be like moving compute into the memory, which would potentially remove the need to move the data in and out of storage as much also.
It feels like computing hardware infrastructure is fundamentally evolving.
In the past platforms have integrated ML algorithms into relational databases and SQL through extensions (both commercial and open source). A famous open source one was MADlib [1], which has an implementation of neural networks. Even the commercial ones were similar, I used ML algorithms in SQL Server many years ago around 2005 I think.
I am wondering about.. SQL as a declarative structured query language that can be optimized into essentially any kind of distributed, directed acyclic graph of processing - vs the special characteristics of relational databases (relational algebra, relvars, etc. etc.) is an important distinction as - of yet, I see the author linking both together so I'm trying to understand what it is about relational structures that specifically helped here (just not seeing it yet, not that it isn't there).
Also, wondering if ISO/IEC 9075-15:2023 SQL standard for multidimensional arrays (MDA) is of any use here? Paper describing SQL/MDA here [2].
Yeah, I mean - I know that people do this in real life, but when I encountered this recently while trying GPT-Live, it was truly annoying - I felt like it kept causing me to lose my train of thought. Perhaps if it was a bit more adaptive or could even be turned off.. but if it does this every.. single... time... I am saying something and I pause for a second... ugh.
Might be interested in orthogonal reading - "The Textual Warehouse" (ISBN-10: 163462954X) by data warehouse pioneer Bill Inmon. He is and always has been ahead of his time with his thinking!
You're assuming that you'll still be able to get a personal computer in the future. With the rate that we are losing the ability to purchase new affordable equipment, I am not sure how much longer personal computers will remain a thing except for hobbyists, and if they will get much more expensive.
Not OP and just guessing, but probably SXM2 GPU modules for the V100. Those can be acquired fairly inexpensively, but there is work to do to get them working together and the V100 has some limitations on the types of models you can run.
I second the idea to use a zip. In fact it's what a lot of vendors do because it is so ubiquitous, even Microsoft for example - the "open Microsoft Office XML document format" is just a zip file containing a bunch of folders and XML files.
- you're on HN, so you have an opportunity to tell us the name of the thing - marketing
- perhaps consider your marketing budget to be the area you need to invest in now, and indirectly how that budget can actually generate a little revenue and exercise the engine
- e.g. do you have packages you can ship through your service to announce it to others? Use your marketing budget to do it and collect your marketplace fees - marketing again. Nothing says I believe in my product like using it for real (IMHO).
- a bold and risky move (?) - most would say fake it till you make it - but I for one tire of this tactic. How about approaching this with honesty and reward the first adopters? "We're brand new but you can be the first to help us prove out this model?".
- your marketplace is a network. Search for an opening that has viral properties. Try to tap into something that has a network effect, go to where your customer is and see how you can target/advertise strategically and respectfully - become a trusted partner to one or more communities (e.g. thinking out loud, eBay sellers maybe?). This could include finding the right partner(s) who have a problem and are willing to give you a shot in an existing network.
On the last point - as an example of the viral thing - I worked on a real estate tool a while back. We found a viral hook - there were for sure properties that needed to be processed and worked through the tool - we email invited the parties involved in the transaction to invite them to work on the property in the secure tool, and we gave them the ability to invite others working the same transaction to the tool, and we focused everything on polishing that workflow and experience.
This way, as soon as one person used the tool they could invite others to use it legitimately to work in the tool and that was the viral aspect.
This points to, replace real estate tool and house with "package" and invite... and can you achieve something viral that spreads itself... like, when someone ships with your tool, it emails the recipient with the link to your site for status tracking and a call to action to make them want to ship using your platform.
This to me is a lot of marketing and product strategy around incentivizing the network effect.
Disclaimer: I never made millions of dollars off a marketplace. But I did help stand up the real estate mechanism I mentioned and that business reliably brought in 5-10 grand a month with no marketing effort and just that one mechanism, and it also helped us find a few key network partners. That's what drives my feedback.
I have a bunch of TI-99 hardware in storage, have been thinking to donate it to a computer museum potentially. I had one in my hand when I was 5 thanks to my grandpa (it made me what I am today!).
I still don't understand what happened to using Apache Avro [1] for row-oriented fast write use cases.
I think by now a lot of people know you can write to Avro and compact to Parquet, and that is a key area of development. I'm not sure of a great solution yet.
Apache Iceberg tables can sit on top of Avro files as one of the storage engines/formats, in addition to Parquet or even the old ORC format.
Apache Hudi[2] was looking into HTAP capabilities - writing in row store, and compacting or merge on read into column store in the background so you can get the best of both worlds. I don't know where they've ended up.
Also, it was paid for by US taxpayer dollars - the entire content should have been released somewhere for free, maybe even someone would have started up a new project to maintain it, for example, something under Wikimedia or some other nonprofit.
This wholesale elimination of valuable information and data owned by the public is so incredibly sad and damaging to our future.
Maybe we need a FOIA request to get the entire contents released to the public.
Half serious - but is that really so different than many apps written by humans?
I've worked on "legacy systems" written 30 to 45 years ago (or more) and still running today (things like green-screen apps written in Pick/Basic, Cobol, etc.). Some of them were written once and subsystems replaced, but some of it is original code.
In systems written in the last.. say, 10 to 20 years, I've seen them undergo drastic rates of change, sometimes full rewrites every few years. This seemed to go hand-in-hand with the rise of agile development (not condemning nor approving of it) - where rapid rates of change were expected.. and often the tech the system was written in was changing rapidly also.
In hardware engineering, I personally also saw a huge move to more frequent design and implementation refreshes to prevent obsolescence issues (some might say this is "planned obsolescence" but it also is done for valid reasons as well).
I think not reading the code anymore TODAY may be a bit premature, but I don't think it's impossible to consider that someday in the nearer than further future, we might be at a point where generative systems have more predictability and maybe even get certified for safety/etc. of the generated code.. leading to truly not reading the code.
I'm not sure it's a good future, or that it's tomorrow, but it might not be beyond the next 20 year timeframe either, it might be sooner.
I like Fastmail with my own domain for personal email, but the reality is nothing is a complete replacement for a Google account, given how tied in it is with auth and the whole Google ecosystem. I still have to use Google for work.
Proton is another one people often suggest. Hey.com sometimes too. No experience with those myself.
There are other options (such as the big guys, iCloud mail or Outlook.com), but aside from self-hosting (which I don't want to spend time maintaining just for my personal mail), I personally haven't seen much outside of those ones that are recommended often.
Not selling anything, but I am trying to figure out what to do to help support solo and micro entrepreneurs, very small businesses (2-3 people) and very small nonprofits.
I feel like there are a lot more people in this position now (me included), but I don't want to do things for the sake of doing them... I want to find out what solo folks really benefit from and help make sure you get more support.
I agree that spending time on inference or compute every time for the same LLM task is wasteful and the results would be less desirable.
But I don't think the thought experiment should end with that. We can continue to engineer and problem solve the shortcomings of the approach, IMHO.
You provided a good example of an optimization - tool creation.
Trying to keep my mind maximially open - one could think of a "design time" performance at runtime - where the user interacting with the system is describing what they want the first time, and the system is assembling the tool (much like we do now with AI assisted coding, but perhaps without even seeing the code).
Once that piece of the system is working it is persisted so no more inference is required, as essentially code - a tool, that saves time. I am thinking of this as essentially memoizing a function body- i.e. generating and persisting the code.
There could even be some process overseeing the generated code/tool to make sure the quality meets some standard and providing automated iteration, testing, etc if needed.
A big problem is if the LLM never converges to the "right" solution on it's on (e.g. the right tool to generate the HTML from the SQL query, without any hallucination). But, I am willing to momentarily punt on that problem as being more to do with the determinism problem and the quality of the result. The issue isn't per se the non-deterministic results of an LLM anyway, it's the quality of the result fit for purpose for the use case.
I think it's difficult but possible to go further with the thought experiment. A system that "builds itself" at runtime, but persists what it builds, based on user interaction and prompting when the result is satisfactory...
I remember one of the first computer science things I learned- the program that could print out it's own source code. Even then we were believing that systems could build themselves and grow themselves.
So my ask would be to look beyond the initial challenge of the first time costs of generating the tool/code and solve that by persisting a suitable result.
What challenge or problem comes next in this idea?
E-mail: sixdimensional ATAT [dimensionsix DOT com] minus the ATAT and brackets.