I used Apache Drill to query json line log files stored in Azure Blob. It is very easy to configure and run it. I used in embedded mode for not so big queries and some visualisation in Apache Superset. It worked really well. I created some views in parquet to speed it up.
Be aware, there is no such thing as schemaless, Drill is schema on-read and if your files contain changing schemas, it is painful to workaround all the errors you face. JSON is too ambiguous when it comes to types.
HelloFresh | Berlin, Germany | Visa | Full-time | Onsite
HelloFresh is the leading global provider of fresh food at home in its 10 markets. It is the biggest meal kit service in the US.
HelloFresh is looking for Data Scientists and Machine Learning Engineers to join the team. We aim to optimise our whole supply chain (from procurement to delivery) using data science and automatic decision making.
It is really good at performance (thanks to optimised C++ implementation) for running it on large networks compared to networkx or other pure python implementations but its usability is very bad.
Unfortunately, Windows is not supported. It relies on systeminformation* package and it does not cover all resources for Windows. PRs are welcome to fix the issues for Windows.
Be aware, there is no such thing as schemaless, Drill is schema on-read and if your files contain changing schemas, it is painful to workaround all the errors you face. JSON is too ambiguous when it comes to types.