That is a sharp and slightly chilling analogy. If Steve Jobs saw the computer as a tool that amplified human effort (the bicycle), and AI represents a tool that automates that effort entirely (the car), then the "obesity epidemic" of the mind is likely Cognitive Atrophy.
The technique can be applied by any engine, not just DataFusion. Each engine would have to know about the indexes in order to make use of them, but the fallback to parquet standard defaults means that the data is still readable by all.
Hilbert curves are used in modern data lakehouse storage optimisation techniques, such as Databricks' "liquid clustering" [1]. This can replace the need for more traditional "hive-style" partitioning, in which the data files are partitioned based on a folder structure (e.g. `mydatafiles/YYYY/MM/DD/`).
I wanted to build a Windows container image but I really did not want to install Docker Desktop. After some digging around I found my way to the Docker server / client binaries for Windows page that allowed me to do this: https://docs.docker.com/engine/install/binaries/#install-ser...
This is interesting... possibly a move by Databricks to try and build on their "data lakehouse" concept to counter the recent "Fabric platform" announcements at MS Build.
Databricks coined the "Delta lake" concept and are still (just about) leading the way, but Fabric has the potential from MS to take away that marketshare. Databricks need to improve their "serverless SQL" offering, and add a serious "data warehouse" component alongside the lake.
Windows 11 is the perfect time to switch to Linux (especially for first-timers), using WSL2. Then, once you're familiar with Linux, it's a shorter jump to replace Windows completely (and appreciate how fast Linux can be without WSL)
- Gemini