Location: Europe, Role Type: Fractional, Remote / Worldwide
Remote: Yes
Willing to relocate: Yes
Technologies: Kafka, Flink, Spark, Cassandra, Scala, Rust / PyO3 / Tokio, Python, Zookeeper, Databricks, Delta Lake, BigQuery, Redshift, Hive, Kinesis, Airbyte, Airflow, Temporal, DBT, Aerospike, Snowflake, PrestoSQL, Trino, Clickhouse, SQL, Vector DBs, Golang, gRPC/protobuf, Terraform, CUDA
Resume: https://drive.google.com/file/d/1lZuSpneLDVzzYoeaA8or-m5m2dfYAMB4/view
Email: in the profile (please do mention "Via HN" in the subject for your email to land in the right place in my inbox)
[Remote Worldwide, Fractional] I'm pursuing a role in the high-scalability distributed systems space. I'm a well-rounded Scala/Rust/Python dev, well-versed in data engineering, with deep knowledge of the internals of distributed datastores. I have experience with data modeling for high throughput database activity and a strong understanding of which workloads and data access patterns are scalable and what datastore the data should reside in. I have a highly confident ability to lead a data / MLops project from start to finish. Hire me as fractional data engineer.
Core Focus Areas:
● Cassandra (Data Modeling, Troubleshooting Performance And Operational Issues)
At some point, I needed to write a function which, given a collection of product titles, picks one that is neither the longest nor the shortest, it should pick the one which best captures the essence of the product while not being excessively verbose. For example, given the product titles below, it should pick "Portable Two-Way Translator, Handheld".
Based on previous experience with centroid-based algos, the function I wrote does a first pass throwing all words from all product titles into one big bag, then computing a centroid (frequency histogram with low-frequency words removed). The second pass is to compute a cosine similarity score for each product title (its own frequency histogram against the centroid). Whichever product title is the most similar to the centroid wins.
That algo may have existed already in some academic paper somewhere, but I came up with it independently.
SEEKING WORK, Senior Data Engineer, Remote / Worldwide
Well-rounded Scala/Rust/Python dev, well-versed in data engineering, with deep knowledge of the internals of distributed datastores. I have experience with data modeling for high throughput database activity and a strong understanding of which workloads and data access patterns are scalable and what datastore the data should reside in. I have a highly confident ability to lead a data / MLops project from start to finish. Hire me.
Core Skills:
● Cassandra (Data Modeling, Troubleshooting Performance And Operational Issues)
I'm not sure if I should be applying but before I do, I'm wondering what kind of novel or interesting data driven algorithms you refer to here. Are you referring to like adaptive query planners based on runtime statistics? Real-time detection of query latency spikes? Automated creation of secondary indexes based on usage patterns? Statistical smarts to decide what to cache and when? These would be sort of runofthemill for any modern database product, can you clarify if you are looking for any of these?
Don't rely on gut feeling or anecdotal evidence, there's good data you can look at.
Over a few years, I came to rely on proxy indicators that I found to be reliably in sync with the strength / weakness of the job market for software engineers.
One of these indicators is the business spending index (it has an official name, something about purchasing managers' sentiment index or similar). CNBC tends to show it a lot these days, you can't miss it.
Another one is "US Auto Loans Delinquent by 90 or More Days (I:USALD90)" which is a very good "finger to the wind" for how the overall economy is doing, since everyone needs a car.
Neither of these indicators are looking good these days.
I use it to extract either vocabulary or grammar patterns from snippets of text written in a foreign language I'm learning. I double-check all output, just as you do.
Looking at this, it's not long until OpenAI starts offering a version of this to ecommerce companies for their search backends, it's a space that is held back by corporate corruption, and it's ripe for disruption.
I'm still working through legal complications with my current employer, and in no mood to start my own company. The job market is also bad on the dataeng / datainfra / mlops side of things. My first guess is that interest rates still have to come down some. Giving up on software engineering and trying a different line of work is not an option for me, I was born with a passion for this.
> Is your company like this? If you have real info and not just suspicions, let’s name some names.
A certain company (the name starts with "A") is widely known for doing this, you might guess which company I refer to by deep-diving into my comments history.
But they too will behave in the same way at some point, due to compliance requirements. Banks and fintech companies need to know 1. who you are 2. that your income is from legit sources.
For #2, ask your customers (who paid you recently) to give you a copy of the banking app records of those payments, where the last 4 digits of the payer's bank account can be seen. Wise can then cross-reference those against what they see on their end, at which point they can declare your source of income verified. (I had to learn this the hard way myself.) Very importantly, those payment records have to be from the banking app that the accounts payable team is using, not Ripple or any internal payroll system.
The custom search engine space is absolutely ripe for disruption.
Whenever you need to add a search experience to your SaaS project, you don't really think of building it yourself, however the established companies in the search engine space are lazy, retrograde and corrupt, and they absolutely don't deserve your money. (I know because I work for one of them.)
OpenAI has recently announced it will slash the cost of calls to its embeddings API by a whopping 75%. This is a huge opportunity for a startup to disrupt the search engine space.
With access to OpenAI's now-affordable embeddings API and to a vector database, one mid-level engineer can build a highly scalable custom search engine for ecommerce and retail, in a few weeks.
We are not at all in disagreement. In fact, you are making the same point I'm making, which is that there is nothing special about any of the commercially available search engines today. With access to OpenAI's embeddings API and to a vector database, a mid-level engineer can build a highly scalable search engine in a few weeks. As a startup, it makes sense today to build your own search engine rather than buy off the shelf.
The only thing companies like ElasticSearch and Algolia still have is their pre-existing customer bases, a few thin layers of marketing, and some network effects. Search engine companies are effectively marketing companies nowadays.
That explains why the corrupt management at my company thinks it's a better strategy to throw hissy fits on social media in the general direction of OpenAI, than to work on actually building useful tech.
[Remote Worldwide, Fractional] I'm pursuing a role in the high-scalability distributed systems space. I'm a well-rounded Scala/Rust/Python dev, well-versed in data engineering, with deep knowledge of the internals of distributed datastores. I have experience with data modeling for high throughput database activity and a strong understanding of which workloads and data access patterns are scalable and what datastore the data should reside in. I have a highly confident ability to lead a data / MLops project from start to finish. Hire me as fractional data engineer.
Core Focus Areas:
● Cassandra (Data Modeling, Troubleshooting Performance And Operational Issues)
● Apache Iceberg (Scaling, Tuning, Self-Hosted Setup)
● Stream Processing At Scale: Kafka, Flink, Spark Streaming, Storm
● Custom-Crafted Contextualized Embeddings, Vector-Based Semantic Search, Deep Intent Recognition In Search Engine Queries
● Languages: Scala, Rust, Python, SQL (proficient), Golang (ramping up)
Educational Background: Computer Science.
Solid experience working remotely and working with teams that are distributed geographically. I typically work Pacific Time hours.
Ask: $385K base, pro-rated.