US consumer spending dashboard built on data from 50M+ cards now on Snowflake(app.snowflake.com)
app.snowflake.com
US consumer spending dashboard built on data from 50M+ cards now on Snowflake
https://app.snowflake.com/marketplace/listing/GZTSZ290BUX40/
85 comments
Getting downvoted for a true statement.
Reading up on programmatic advertising from a perspective of learning it vs resources that critique it is illuminating.
You’re describing how it works. Use cash and buy offline with your cellphone left at home (yes, all 3), or get your data sold for other people’s products.
I used to be more vocal about privacy but reviewing the innards of the tech make me think this dynamic will never change. Too much money and too well built.
Reading up on programmatic advertising from a perspective of learning it vs resources that critique it is illuminating.
You’re describing how it works. Use cash and buy offline with your cellphone left at home (yes, all 3), or get your data sold for other people’s products.
I used to be more vocal about privacy but reviewing the innards of the tech make me think this dynamic will never change. Too much money and too well built.
I'm interested in taking this approach to learning, would you mind sharing some good starting points?
Dominik Kosorin books on programmatic advertising on Amazon are a great start. First place I’ve seen Facebook and Twitter referred to only as “advertising platforms,” and the candor was nice.
> with your cellphone left at home
What's the angle on this?
What's the angle on this?
Dozens of companies buy location data from mobile apps, ad networks, cell networks, wifi networks, etc. and resell it. Some of the platforms will let you geofence your store then track phones back to where they "go home" to charge that night and then mail coupons to that address.
https://heyirys.com/ https://www.redmob.io/ https://echo-analytics.com/ https://www.factori.ai/ among others.
https://heyirys.com/ https://www.redmob.io/ https://echo-analytics.com/ https://www.factori.ai/ among others.
I’m loose on the terminology but there’s a market for “offline data” correlation.
Gist of it is if your cellphone says you went wherever, there’re intuitions about what you bought. For example, stores have Bluetooth and Wi-Fi beacons on top of aisles - if your phone beacons but doesn’t connect, the store can pull a device ID (so to speak), and again correlate.
Similarly interesting is “retargeting” that tracks you’ve unsubscribed from emails and the retargets a similar ad.
The Dominik Kosorin books on adtech for a technical audience (Amazon) were illuminating and also broke my spirt on privacy. This stuff is everywhere and there’s an identity graph out there that has your life. No dodging it until regulations change or you go offline.
Edit - some other interesting angles. Third party cookie blocks in browsers harmed the industry. One of the main suggested recovery approaches is building graphs on authenticated logins. Every resource accessed post-login has a good chance of being that user. This is behind all the “sign in with Google/FB,” and known for a while as part of adtech. The primary issue with authn’d tracking is the friction of logging in, and the power it gives to closed networks that can track post-login easily like Google, FB, Apple.
What is the new, de-frictioned login flow pushed widely now as a way to go passwordless and solve the many security problems? All do-good messaging? Passkeys. What is also a de-fricituoned post-authentication user graph, if you use it to sign into everything? Passkeys. Adtech is a cancerous industry.
Gist of it is if your cellphone says you went wherever, there’re intuitions about what you bought. For example, stores have Bluetooth and Wi-Fi beacons on top of aisles - if your phone beacons but doesn’t connect, the store can pull a device ID (so to speak), and again correlate.
Similarly interesting is “retargeting” that tracks you’ve unsubscribed from emails and the retargets a similar ad.
The Dominik Kosorin books on adtech for a technical audience (Amazon) were illuminating and also broke my spirt on privacy. This stuff is everywhere and there’s an identity graph out there that has your life. No dodging it until regulations change or you go offline.
Edit - some other interesting angles. Third party cookie blocks in browsers harmed the industry. One of the main suggested recovery approaches is building graphs on authenticated logins. Every resource accessed post-login has a good chance of being that user. This is behind all the “sign in with Google/FB,” and known for a while as part of adtech. The primary issue with authn’d tracking is the friction of logging in, and the power it gives to closed networks that can track post-login easily like Google, FB, Apple.
What is the new, de-frictioned login flow pushed widely now as a way to go passwordless and solve the many security problems? All do-good messaging? Passkeys. What is also a de-fricituoned post-authentication user graph, if you use it to sign into everything? Passkeys. Adtech is a cancerous industry.
So there isn’t a record of your phone going to the store…
Oh but you get magic points that make up for everything. Which you already pay for via outsided transaction fees every time you use your card. And predatory interest rates if you happen to miss a single payment. And get your personal info and spending habits sold on top of that. The entire industry is a scam.
Giving bazookas to unscreened people.
Hard to tell if this is interesting without seeing a sample of the data. Though the data originates from "50M+" cards, the summary doesn't make clear how many rows of data this is. Might not be that many rows since it doesn't get any finer than weekly granularity.
Hey founder here -
It would be indeed great if Snowflake allowed us to preview the data. That said, if you have a Snowflake account, you can mount it and automatically get a trial (you can run arbitrary queries against it).
The data is aggregated at a weekly level, by category, merchant, and demographics. It is not single individuals' data if that is what you are after.
It would be indeed great if Snowflake allowed us to preview the data. That said, if you have a Snowflake account, you can mount it and automatically get a trial (you can run arbitrary queries against it).
The data is aggregated at a weekly level, by category, merchant, and demographics. It is not single individuals' data if that is what you are after.
Interesting product - compared to competitors like SecondMeasure, what differentiates your panel/data cleaning approach vs. everyone else (since I assume the data is sourced from the same place as you competitors)?
I am a big admirer of what Second Measure built before their acquisition.
- Different data sources - More accurate (though this is of course debateable, our plan is to publish benchmarks, accuracy, transparency, etc.) but this will always be debateable - Focus on data scientists & Snowflake users (rather than a SaaS platform)
Obviously, this is a very early product. The key is to join types of datasets together, while maintaining accuracy (see vision outlined here: https://magis.substack.com/p/datanomics)
- Different data sources - More accurate (though this is of course debateable, our plan is to publish benchmarks, accuracy, transparency, etc.) but this will always be debateable - Focus on data scientists & Snowflake users (rather than a SaaS platform)
Obviously, this is a very early product. The key is to join types of datasets together, while maintaining accuracy (see vision outlined here: https://magis.substack.com/p/datanomics)
How is this different that what IRI sells?
Or Argus?
Or the various merchant exchanges (e.g. WMX/Walmart, 8451/Kroger)?
Genuinely curious.
Or Argus?
Or the various merchant exchanges (e.g. WMX/Walmart, 8451/Kroger)?
Genuinely curious.
- IRI sells syndicate (aggregate) data on CPG sales -- individual brands in the store -- not at the merchant level (ie. Chipotle)
- Kroger/Walmart sell more detailed versions of that, but for their own stores AFAIK
- I am not familiar enough w/ Argus to know to be honest, but I do not see them come up for competitive intelligence use cases (which is what we are aiming for) outside of card issuers/merchant acquirers themselves.
That said, we obviously have competitors. We are trying to differentiate with our data accuracy, focus on data scientists, and focus on Snowflake approach.
Obviously, this is a first product: we hope to add to it. See the vision here: https://magis.substack.com/p/datanomics
That said, we obviously have competitors. We are trying to differentiate with our data accuracy, focus on data scientists, and focus on Snowflake approach.
Obviously, this is a first product: we hope to add to it. See the vision here: https://magis.substack.com/p/datanomics
Thanks.
So what’s your actual business name and marketing site for me to read more?
So what’s your actual business name and marketing site for me to read more?
You keep saying that in this thread, founder of what? Snowflake?
Cybersyn
I can't seem to find you on the list of registered data brokers in the State of California [1]. I'm submitting a complaint to flag it to them now.
1. https://oag.ca.gov/data-brokers
1. https://oag.ca.gov/data-brokers
Founder of what?
It is normal to state that and to be explicit
Following the link in his bio you'll find his LinkedIn and other documentation pointing to him being a founder at Cybersyn, so it does look like he can speak for this, but it isn't surprising that we got such a vague answer since honesty would require admitting how shady this is.
Thank you. My first thought now is that this is why we have GDPR and why it is so important.
The docs are more helpful then the link above:
https://docs.cybersyn.com/our-data-products/consumer/consume...
https://docs.cybersyn.com/our-data-products/consumer/consume...
It seems odd that you can't even preview a small cached subset (10 rows?) of a table in Snowflake. I don't even remember a time I couldn't see a table preview in BigQuery, so surely it's not a technical limitation?
I definitely agree Snowflake Marketplace should have this.. there is no technical limitation.
They could have baked it in to the descriptions or code samples.
The assumption - and maybe it didn't work - was that folks would just try the data (it is free to try). I do think one great thing about the Snowflake Marketplace is how the trials work -- you get a lot of flexibility. Of course, you need a capacity Snowflake account to do so -- but the product is meant for enterprise users.
with all of the fintech apps that attach to a card, i wonder if they break down spending categories by vice. how many people paying their dealer or other vice related friends with cashapp, venmo, etc, and then bigData making the connection of what the transactions really were?
Hey there - founder here. We can indeed break down by payments in these FinTech apps. There are privacy concerns with certain aspects of this, which we of course could not touch. However, measuring the "shadow" economy, of course, serves a legitimate purpose.
you're not doing yourself any favors with the privacy minded folks that think people analyzing the data they did not provide to you directly for that purpose is something to be removed like cancer. but hey, at least you're honest about the granularity of the unintended data you can discern from the data you've purchased about everyday people.
I wasn't quite clear - we don't have access to identifying information in the data (ie. we could measure Venmo but not that John Doe paid $X to Jane Doe at phone number XXX-XXXX)
I look at my statements, and each transaction from an app adds enough information to the memo line that I can ID the other user. I’m assuming the data you’ve purchased is “anonymized”, so I don’t know what that line item looks to you. But you still have time stamps and app names. Are you telling me you honestly could not deanonymize. Or are you telling me that you’re just not at this time?
> analyzing the data they did not provide to you directly for that purpose is something to be removed like cancer
For the record, cancer is an appropriate metaphor to describe the data broker industry
For the record, cancer is an appropriate metaphor to describe the data broker industry
Talking to adtech about their products vs criticizing their products is the most illuminating thing a privacy-minded person could do.
The nonsense “we care about privacy” wrappers go away and they start talking product-speak about the details. I don’t think privacy regulations change until adtech product leads get comfortable enough talking publicly about their deadpan views on what they can find out about users.
The nonsense “we care about privacy” wrappers go away and they start talking product-speak about the details. I don’t think privacy regulations change until adtech product leads get comfortable enough talking publicly about their deadpan views on what they can find out about users.
[deleted]
There used to be a very funny website called Vicemo that would scrape and display (public by default!) Venmo transactions containing certain keywords/emojis
https://www.thecut.com/2015/02/vicemo-collects-all-your-sket...
https://www.thecut.com/2015/02/vicemo-collects-all-your-sket...
Yeah interestingly, Venmo makes (or at least used to make) every public transaction available in their API
I’ve seen a lot of people promising to sell (aggregated, cleaned, private, inferred) data on Snowflake — using the database as a platform. Snowflake themselves have been listing this idea as a key value proposition.
This is the first that I’ve seen “in the wild” (i.e. outside of an interview with a start-up that yet had to find a lot of paying customers).
Is that common? Do you know of other companies doing that? Do they typically offer additional individual data (say: lead generation), inferred insights (say: likely spending by post-code based on local shops and disposable income), or something else (custom dashboard)?
This is the first that I’ve seen “in the wild” (i.e. outside of an interview with a start-up that yet had to find a lot of paying customers).
Is that common? Do you know of other companies doing that? Do they typically offer additional individual data (say: lead generation), inferred insights (say: likely spending by post-code based on local shops and disposable income), or something else (custom dashboard)?
Most of the large data providers (S&P, Factset, Nielsen, ZoomInfo) have started to do so. I would say that they have partial solutions -- they do not list their entire catalog of assets on the platform.
Cybersyn is obviously a startup explicitly trying to do just this -- AFAIK, we're the first "Snowflake Native" startup but it is hard to tell, there are more than 2200 data products in Snowflake Marketplace already so I easily could be wrong.
Cybersyn is obviously a startup explicitly trying to do just this -- AFAIK, we're the first "Snowflake Native" startup but it is hard to tell, there are more than 2200 data products in Snowflake Marketplace already so I easily could be wrong.
The top of the free listing is repackaged open-source datasets (economic data, Covid).
The top of the paid listing is the usual suspects: weather, geographic (GeoIP and matching elements), and you, CyberSyn, but nothing really exotic yet.
The top of the paid listing is the usual suspects: weather, geographic (GeoIP and matching elements), and you, CyberSyn, but nothing really exotic yet.
That's true but you would sort of expect those datasets to be the most popular just given the addressable market for those is by far the biggest (essentially every enterprise probably requires some degree of weather, Covid, and economic data).
What sort of exotic data would be useful to you, out of curiosity?
What sort of exotic data would be useful to you, out of curiosity?
I personally think opinionated high-value data would be more helpful, a bit like classifying US voters into seven groups has proven transformational for electoral campaigns.
I’ve seen B2B sales desperately trying to sling their solutions to so many companies that are not at the right stage to value or understand it, so things like “companies with one level of engineering management” vs. “two” vs. “more.” “Companies disillusioned by terrible attempts at Agile,” or “… where the security team is a candle they forgot to light,” or “Companies with a violently dysfunctional Product leadership,” or “Companies with an inexperienced/checked-out CEO” or “…looking for a buyer” might be relevant — assuming you manage to rephrase those into less hurtful categories. Details gained from LinkedIn, etc. would be helpful to train our funnel model to say: “We seem more likely to sell our LLMaaS to companies with a certain number of technical employees or CTOs that like open-source.”
Same for customers: I receive so many coupons for things I clearly do not need or care for. But there are things that I would be considering (my Amazon basket is a good indicator). You’d need to include relevant ideas and exclude things I’ve likely already bought, though… Given how Amazon Prime is unable to not insistently recommend TV series that I’ve finished and will hide what I’m actively watching even when I search for it by name, that could be a hard problem to fix.
Matching that with the political thing: how to convince people to buy certain things. There’s been some controversy in trying to use the OCEAN model to get people to vote for certain people, but I think there’s room for an ethical way of telling marketers: that person would subscribe to HelloFresh because it’s cheaper than GrubHub, that person because they need structure in their life, that person because they want to eat healthy, that person because they like the idea of learning something, that person because they need to have something nice to serve their dates and that person, they already subscribe, stop being weird.
I’ve seen B2B sales desperately trying to sling their solutions to so many companies that are not at the right stage to value or understand it, so things like “companies with one level of engineering management” vs. “two” vs. “more.” “Companies disillusioned by terrible attempts at Agile,” or “… where the security team is a candle they forgot to light,” or “Companies with a violently dysfunctional Product leadership,” or “Companies with an inexperienced/checked-out CEO” or “…looking for a buyer” might be relevant — assuming you manage to rephrase those into less hurtful categories. Details gained from LinkedIn, etc. would be helpful to train our funnel model to say: “We seem more likely to sell our LLMaaS to companies with a certain number of technical employees or CTOs that like open-source.”
Same for customers: I receive so many coupons for things I clearly do not need or care for. But there are things that I would be considering (my Amazon basket is a good indicator). You’d need to include relevant ideas and exclude things I’ve likely already bought, though… Given how Amazon Prime is unable to not insistently recommend TV series that I’ve finished and will hide what I’m actively watching even when I search for it by name, that could be a hard problem to fix.
Matching that with the political thing: how to convince people to buy certain things. There’s been some controversy in trying to use the OCEAN model to get people to vote for certain people, but I think there’s room for an ethical way of telling marketers: that person would subscribe to HelloFresh because it’s cheaper than GrubHub, that person because they need structure in their life, that person because they want to eat healthy, that person because they like the idea of learning something, that person because they need to have something nice to serve their dates and that person, they already subscribe, stop being weird.
I knew I've seen the name "Cybersyn" before:
https://en.wikipedia.org/wiki/Project_Cybersyn
https://en.wikipedia.org/wiki/Project_Cybersyn
This is where the inspiration came from, indeed. https://magis.substack.com/p/project-cybersyn
Who wrote the sample query? If I was trying to "Compare Chipotle’s sales performance to that of McDonald’s", my query would not start with "SELECT *". There would be some meaningful columns.
Yeah that's fair. The tables are all in an EAV format (narrow), so the number of columns coming along here is very small (you would filter for Chipotle/McDonalds in your WHERE) but it is fair that that sample query could be more instructive / better practice to name columns.
I'm curious how cybersyn sources the data (partnership with data sellers? etc?) quality, reliability, and freshness are some that come to mind. Entry point using snowflake marketplace (series a investor) seems reasonable. We're missing the aha moments for companies when they join cybersyn datasets to their own, that story is hopefully in the works.
In general, we're sourcing from a variety of 1st party sources.
But yes, the long term vision is to make this all joinable (https://magis.substack.com/p/datanomics)
But yes, the long term vision is to make this all joinable (https://magis.substack.com/p/datanomics)
Does this dashboard cost 2k/m?..
For those who can use this data to predict something/use it in a model, I guess it's pocket money. This is some very powerful data in the right hands.
To add to that, 2k/mo is pretty much the "go-to" pricing for most B2B SaaS products.
To add to that, 2k/mo is pretty much the "go-to" pricing for most B2B SaaS products.
This is correct. Note you are not getting just the dashboard, but you are getting the underlying data itself too -- so you can write arbitrary queries. That said, it is meant for large enterprises that have a clear path to ROI.
This begs the question: where is the button I can press to see if my personal data is included and a quick and easy way to inform you to remove it from your service?
No personally identifying information in the dataset, so I imagine you would have to find some way to deanonymize yourself from clues in the data.
How is it verified that the data is trustworthy then?
$2k/mo for consumer credit card data is absurdly cheap by at least an order of magnitude. So it’s either mispriced or it’s garbage data.
White screen on iOS… render error?
Might be a render error on the Snowflake side -- check https://docs.cybersyn.com/our-data-products/consumer/consume... instead
what's your tech stack to process this data prior to loading to snowflake?
We're doing almost everything entirely in Snowflake. Snowflake is our lead investor (https://www.reuters.com/technology/data-startup-cybersyn-rai...) and we've found it extremely helpful to build entirely on their technology.
We're ingesting from S3 or FTPs usually.
We're ingesting from S3 or FTPs usually.
Any plans for cheaper plans with less or older data?
Sorry to disappoint but unlikely.
This type of product is generally meant for very large enterprises, so this is already an entry-level product at best.
We do make a lot of related economic data available for free though that is published by government sources. I think often what is publicly available and collected by the government is under-appreciated - the level of granularity of government inflation statistics by product category is outstanding for example.
This type of product is generally meant for very large enterprises, so this is already an entry-level product at best.
We do make a lot of related economic data available for free though that is published by government sources. I think often what is publicly available and collected by the government is under-appreciated - the level of granularity of government inflation statistics by product category is outstanding for example.
Reeks of selection bias.
[deleted]
Sorry a small thread hijack. I may be considering returning to my previous employer who has since I left started using Snowflake. Is that unambiguously good or bad? If it can be either, are there some questions I should check? (As a background, I am quite happy writing my queries in SQL)
As a data warehouse it’s great. It’s not a complete solution for an analytics platform. Snowpark seems are bolted on and doesn’t really benefit from the core engine. It’s reporting capabilities are a bit weak. So you’ll probably want some external BI tools and/or tools like R and Python if you’re heavily into analytics.
“It depends”
There’s nothing inherently wrong with Snowflake. But like any database it has a time and place.
There’s nothing inherently wrong with Snowflake. But like any database it has a time and place.
In my experience it excels at the stated use case of OLAP in the cloud with compute separate from storage and claim it's unambiguously good for this.
For OLTP it would be unambiguously bad (although maybe there's hope with Unistore, which I haven't tried.)
Many workloads are a mix and so it can become ambiguous whether it's a great fit / great value / whatever you're defining good/bad-ness by
For OLTP it would be unambiguously bad (although maybe there's hope with Unistore, which I haven't tried.)
Many workloads are a mix and so it can become ambiguous whether it's a great fit / great value / whatever you're defining good/bad-ness by
We're building our entire company almost exclusively on the Snowflake stack. The fact that that is possible, definitely shows you how Snowflake has evolved from just a "warehouse" to a Data Platform (or Data Cloud to use their terminology).
The Marketplace, Streamlit, and Native Apps stand out as particularly cool/useful.
More than anything, just the completeness of the platform (ie. data quality tools, governance, etc.) is super helpful to not have to cobble together.
The AI features they announced today also seem game-changing.
Of course, full diclosure, they are my lead investor but I chose to work with them for these reasons.
The Marketplace, Streamlit, and Native Apps stand out as particularly cool/useful.
More than anything, just the completeness of the platform (ie. data quality tools, governance, etc.) is super helpful to not have to cobble together.
The AI features they announced today also seem game-changing.
Of course, full diclosure, they are my lead investor but I chose to work with them for these reasons.
I'm quite interested in your consumer spending data. Are you able to share what the underlying source/partner is?
Also - wanted to ask about the nominal value projections for market share. Do you some how normalize or weight/calibrate this data? Thanks
Shoot me at email at [email protected]
[deleted]
Users likely not reimbursed for their data getting packaged and resold by merchants, payment networks, or credit card issuers