Nearly 20% of active Twitter accounts likely to be fake or spam(sparktoro.com)
sparktoro.com
Nearly 20% of active Twitter accounts likely to be fake or spam
https://sparktoro.com/blog/sparktoro-followerwonk-joint-twitter-analysis-19-42-of-active-accounts-are-fake-or-spam/
394 comments
Parag apparently lost his patience with superficial and misleading claims about Twitter spam (like this analysis) and posted about it today.
You can see it here (https://twitter.com/paraga/status/1526237578843672576).
Noteworthy highlights:
* Twitter estimates its <5% number from human analysis of multi-thousand user random samplings of mDAU
* Twitter allows that number to remain so high to avoid introducing friction like captcha into real users' experiences
* Twitter uses all sorts of internal private data in its analysis
* Parag says you cannot get a reliable indication of bot/not bot without this internal private data
Having just finished building a Twitter analysis tool, I agree with Parag that the Twitter API doesn't provide sufficient clarity to make decisions about spam. This article's analysis doesn't hold up - just because you can name several features you're going to use to generate a spam confidence score about an account does not mean that spam confidence score will have any precision.
You can see it here (https://twitter.com/paraga/status/1526237578843672576).
Noteworthy highlights:
* Twitter estimates its <5% number from human analysis of multi-thousand user random samplings of mDAU
* Twitter allows that number to remain so high to avoid introducing friction like captcha into real users' experiences
* Twitter uses all sorts of internal private data in its analysis
* Parag says you cannot get a reliable indication of bot/not bot without this internal private data
Having just finished building a Twitter analysis tool, I agree with Parag that the Twitter API doesn't provide sufficient clarity to make decisions about spam. This article's analysis doesn't hold up - just because you can name several features you're going to use to generate a spam confidence score about an account does not mean that spam confidence score will have any precision.
"We undercount active users whose accounts are protected, accounts that view tweets but don’t send any, and accounts that log in and engage in other ways beyond tweeting (like favoriting or adding profiles to lists)."
My markup. If I understand correctly, not having a public tweet is a marker for being a spam account. Isn't that kind of a lot of people? I know from other forums that there's a large ratio between lurkers and active posters.
Given that you need an account to customise your timeline, and, these days, pretty much for just reading a tweet, there may be loads of real reader accounts that never post and never bother customizing their profile.
My markup. If I understand correctly, not having a public tweet is a marker for being a spam account. Isn't that kind of a lot of people? I know from other forums that there's a large ratio between lurkers and active posters.
Given that you need an account to customise your timeline, and, these days, pretty much for just reading a tweet, there may be loads of real reader accounts that never post and never bother customizing their profile.
Oh, do I have notes on their methodology.
1) They talk about "active" accounts (meaning have tweeted in the last 9 weeks), and do a bunch of filtering against that. That seems like a huge bias - lurkers exist, and in my experience are usually the majority of users...this step removes them or ignores them entirely. Frankly, until recently, my twitter account would have been one of the ones they would have discarded as inactive. This one thing alone makes me question all of the rest of their results.
2) By the same token, the rate or frequency with which a user sends tweets has no relation to whether a user is monetizable. If they're seeing ads, they're monetizable...lurkers are just as monetizable as high-volume posters.
1) They talk about "active" accounts (meaning have tweeted in the last 9 weeks), and do a bunch of filtering against that. That seems like a huge bias - lurkers exist, and in my experience are usually the majority of users...this step removes them or ignores them entirely. Frankly, until recently, my twitter account would have been one of the ones they would have discarded as inactive. This one thing alone makes me question all of the rest of their results.
2) By the same token, the rate or frequency with which a user sends tweets has no relation to whether a user is monetizable. If they're seeing ads, they're monetizable...lurkers are just as monetizable as high-volume posters.
This study is a great example of how you can use the data you have available to talk yourself into your conclusions. The implicit point of the study is to refute the "5% of Twitter accounts are spam" stat from Twitter's 10-Q that was the basis of his putting the twitter acquisition "on hold".
Except - the baseline that they choose is entirely NOT comparable to that of Twitter's baseline. The study says:
> Followerwonk selected a random sample from only those accounts that had public tweets published to their profile in the last 90 days, a clear indication of “activity.” Further, Followerwonk regularly updates its profile database (every 30 days) to remove any protected or deleted accounts. We believe this sample is both large enough in size to be statistically significant, and curated to most closely resemble what Twitter might consider a monetizable Daily Active User (mDAU).
Except that we know what Twitter defines as a monetizable DAU:
> We define monetizable daily active usage or users (mDAU) as Twitter users who logged in and accessed Twitter on any given day through Twitter.com or Twitter applications that are able to show ads.
Nothing about posting, nothing about engagement at all - simply: were you able to see an ad?
So there isn't any reason to claim that this "might" represent what Twitter uses as an mDAU - we know, in fact, that is not how they measure it. A more honest statement would have been:
"We selected a random sample of (etc. etc.). We believe that this sample is large enough to be significantly significant, however, it can not be compared to Twitter's mDAU set, as it does not count passive consumers of Twitter content. Instead, this data can be used to suggest that a significant amount of the total posted content on Twitter is delivered by bots"
My guess is the number of consumers of content is greater than the posters of content by several orders of magnitude, though some of that would be mitigated by the longer time horizon.
Except - the baseline that they choose is entirely NOT comparable to that of Twitter's baseline. The study says:
> Followerwonk selected a random sample from only those accounts that had public tweets published to their profile in the last 90 days, a clear indication of “activity.” Further, Followerwonk regularly updates its profile database (every 30 days) to remove any protected or deleted accounts. We believe this sample is both large enough in size to be statistically significant, and curated to most closely resemble what Twitter might consider a monetizable Daily Active User (mDAU).
Except that we know what Twitter defines as a monetizable DAU:
> We define monetizable daily active usage or users (mDAU) as Twitter users who logged in and accessed Twitter on any given day through Twitter.com or Twitter applications that are able to show ads.
Nothing about posting, nothing about engagement at all - simply: were you able to see an ad?
So there isn't any reason to claim that this "might" represent what Twitter uses as an mDAU - we know, in fact, that is not how they measure it. A more honest statement would have been:
"We selected a random sample of (etc. etc.). We believe that this sample is large enough to be significantly significant, however, it can not be compared to Twitter's mDAU set, as it does not count passive consumers of Twitter content. Instead, this data can be used to suggest that a significant amount of the total posted content on Twitter is delivered by bots"
My guess is the number of consumers of content is greater than the posters of content by several orders of magnitude, though some of that would be mitigated by the longer time horizon.
You only need to see who the author of this post is to know that the methodology is crap, the numbers are likely made up (19.42% is WAY too specific), and the post is just a grab for media attention on the coattails of some other internet meme garbage.
This guy (Rand Fishkin) has been selling SEO as a religion for the better part of this century, and is in no small part responsible for all the search-result-garbage style websites everyone is complaining about elsewhere on HN today and every other day.
He's a third-rate market-bro hack that's been taking advantage of web professionals who get thrown into SEO/Marketing jobs and have no idea what they're doing by relentlessly shoving half-assed corporate strategies through moz.com and now his new sparktoro.com, and calling himself the great SEO redeemer.
Wanna question his methodology? There is none. Wanna question his science? Totally devoid.
This guy (Rand Fishkin) has been selling SEO as a religion for the better part of this century, and is in no small part responsible for all the search-result-garbage style websites everyone is complaining about elsewhere on HN today and every other day.
He's a third-rate market-bro hack that's been taking advantage of web professionals who get thrown into SEO/Marketing jobs and have no idea what they're doing by relentlessly shoving half-assed corporate strategies through moz.com and now his new sparktoro.com, and calling himself the great SEO redeemer.
Wanna question his methodology? There is none. Wanna question his science? Totally devoid.
I signed up with a vpn and got banned for life after a single nonsense tweet about not liking the feed and following 4-5 famous people. I don’t think bot detection techniques are very robust
> This methodology likely undercounts spam and fake accounts, but almost never includes false positives (i.e. claiming an account is fake when it isn’t).
In other words, their model performs well on their training set, and they don’t acknowledge that it may be over fitted or mislabeled, and they hand wave mistakes
> This methodology likely undercounts spam and fake accounts, but almost never includes false positives (i.e. claiming an account is fake when it isn’t).
In other words, their model performs well on their training set, and they don’t acknowledge that it may be over fitted or mislabeled, and they hand wave mistakes
"70.23% of @ElonMusk followers are unlikely to be authentic, active users who see his tweets."
Somewhat alarming, if accurate.
Somewhat alarming, if accurate.
This doesn't come as a surprise; I was able to buy 1 MILLION fake followers on Instagram. The account is still alive and well.
I've been meaning to write a blog post regarding this endeavor. But moreso, I've come to the conclusion that social media needs to have verifiable audits for their userbase; similar to how there are audits done for financials. A lot of the value of these companies IS derived from their DAU/MAU and or userbase in general (example: WhatsApp - $19 billion for their 1 billion users).
I've been meaning to write a blog post regarding this endeavor. But moreso, I've come to the conclusion that social media needs to have verifiable audits for their userbase; similar to how there are audits done for financials. A lot of the value of these companies IS derived from their DAU/MAU and or userbase in general (example: WhatsApp - $19 billion for their 1 billion users).
I signed up at Twitter 9 years ago and use it every day, but have less than 50 tweets. Either you want to engage in discussion (to a varying degree) or you simply don't. But the latter doesn't necessarily mean you don't observe the discussion.
Of course an active Twitter account can be one that never tweeted in 10 years. Such an account may as well see advertisements etc... So I think any outside studies are fundamentally flawed since they don't have access to internal data like last login time.
Of course an active Twitter account can be one that never tweeted in 10 years. Such an account may as well see advertisements etc... So I think any outside studies are fundamentally flawed since they don't have access to internal data like last login time.
64.56432% of significant digits are misused.
Interesting to think about:
> Our systems do not, however, attempt to identify Twitter accounts that may be irregularly operated by a human but have some automated behaviors (e.g. a company account with multiple users, like our own @SparkToro, or a community account run by a single person, like Aleyda Solis’ @CrawlingMondays). We cannot know how Twitter (or Mr. Musk) might choose to classify these accounts, but we bias to a relatively conservative interpretation of “Spam/Fake.”
So this means:
* @EmojiAquarium - spam/fake * @threateningcake - not spam/fake * @CanYouPetTheDog - spam/fake * @ChuckGrassley - not spam/fake (?? - what fraction are staff generated vs Chuck?) * @Wendys - spam/fake * @Twitter - spam/fake
> Our systems do not, however, attempt to identify Twitter accounts that may be irregularly operated by a human but have some automated behaviors (e.g. a company account with multiple users, like our own @SparkToro, or a community account run by a single person, like Aleyda Solis’ @CrawlingMondays). We cannot know how Twitter (or Mr. Musk) might choose to classify these accounts, but we bias to a relatively conservative interpretation of “Spam/Fake.”
So this means:
* @EmojiAquarium - spam/fake * @threateningcake - not spam/fake * @CanYouPetTheDog - spam/fake * @ChuckGrassley - not spam/fake (?? - what fraction are staff generated vs Chuck?) * @Wendys - spam/fake * @Twitter - spam/fake
To be clear: it doesn't matter, as far as Elon Musk and his buyout is concerned.
To quote Matt Levine's "Money Stuff" newsletter:
> “Temporarily on hold” is not a thing. Elon Musk has signed a binding contract requiring him to buy Twitter.
> That contract does not allow Musk to walk away if it turns out that “spam/fake accounts” represent more than 5% of Twitter users... The merger agreement contains a provision that allows Musk to walk away if Twitter’s securities filings are wrong ... but only if the inaccuracy would have a “Material Adverse Effect” on the company. That is an incredibly high standard: Delaware courts have almost never found an MAE.
> Musk ... had the opportunity to do due diligence on these numbers before signing the deal. (He declined.) He can’t now go to Twitter and say “actually now you need to prove that your user numbers are right.”
[0]https://www.bloomberg.com/opinion/articles/2022-05-13/elon-m...
To quote Matt Levine's "Money Stuff" newsletter:
> “Temporarily on hold” is not a thing. Elon Musk has signed a binding contract requiring him to buy Twitter.
> That contract does not allow Musk to walk away if it turns out that “spam/fake accounts” represent more than 5% of Twitter users... The merger agreement contains a provision that allows Musk to walk away if Twitter’s securities filings are wrong ... but only if the inaccuracy would have a “Material Adverse Effect” on the company. That is an incredibly high standard: Delaware courts have almost never found an MAE.
> Musk ... had the opportunity to do due diligence on these numbers before signing the deal. (He declined.) He can’t now go to Twitter and say “actually now you need to prove that your user numbers are right.”
[0]https://www.bloomberg.com/opinion/articles/2022-05-13/elon-m...
What kind of idiot would attempt a purchase of such magnitude before doing his due diligence (obviously, no one, it's an excuse after the price collapsed)? For fs some people probe fruit for a minute until they decide it's worthy a purchase.
How many more of these types of "events" until people stop treating Elon like some super genius god? Stop giving/loaning him money enabling this garbage.
How many more of these types of "events" until people stop treating Elon like some super genius god? Stop giving/loaning him money enabling this garbage.
I think I've missed something here. Why would Elon not want to buy twitter if it has >5% spam accounts? Isn't the number of fake accounts one of the reasons he wanted to buy it in the first place? I don't understand why this would make him want to back out. It doesn't seem to be in conflict with any of the reasons he wants to buy.
[deleted]
If all you want to know is what fraction of all twitter accounts are spam accounts, it should be really easy:
1. Select 1000 accounts uniformly at random. Either from among all twitter accounts, or from active twitter accounts for whatever definition of "active".
2. Classify these 1000 by hand. Do as much investigation into them as you need to classify them accurately; no need to use heuristics here.
You will (with very high probability) get an estimate accurate to within a percent or so. If you do statistics you could find the actual bounds.
1. Select 1000 accounts uniformly at random. Either from among all twitter accounts, or from active twitter accounts for whatever definition of "active".
2. Classify these 1000 by hand. Do as much investigation into them as you need to classify them accurately; no need to use heuristics here.
You will (with very high probability) get an estimate accurate to within a percent or so. If you do statistics you could find the actual bounds.
Remember when Twitter first started, people thought Twitter usernames were going to be like domain names (except free)? LOL. I must have 50 Twitter accounts personally. Any time I had an idea, I used to grab a Twitter account for it.
Anyway I think it’s pretty goofy to try to make claims around %s of Twitter accounts, active or otherwise, only absolute numbers make sense. What really matters at this point besides revenue and revenue trend?
Anyway I think it’s pretty goofy to try to make claims around %s of Twitter accounts, active or otherwise, only absolute numbers make sense. What really matters at this point besides revenue and revenue trend?
Cool effort! Does seem like a lot of thinking went into this. But, a few points:
"Through trial and error (and, of course, pattern-fitting) we crafted a scoring system that could correctly identify over 65% of the spam accounts."
65% is not actually very accurate for a binary classifier...
"Applying this model to the ~44K random, recently-active accounts provided Followerwonk produces a quality score for each account, visualized below:"
Many real twitter uses are likely not to be "active" aside from reading stuff. So this methodology would clearly overestimate the number of spam/fake accounts (which all would be active).
Also, this is an important point:
"The other potential critique is our spam/fake follower calculation methodology. Because we crafted it in 2018, based off sample sets of purchased spam accounts, it’s likely that more sophisticated spammers and fake accounts go unidentified by our system"
The features collected are certainly outdated by now.
"Through trial and error (and, of course, pattern-fitting) we crafted a scoring system that could correctly identify over 65% of the spam accounts."
65% is not actually very accurate for a binary classifier...
"Applying this model to the ~44K random, recently-active accounts provided Followerwonk produces a quality score for each account, visualized below:"
Many real twitter uses are likely not to be "active" aside from reading stuff. So this methodology would clearly overestimate the number of spam/fake accounts (which all would be active).
Also, this is an important point:
"The other potential critique is our spam/fake follower calculation methodology. Because we crafted it in 2018, based off sample sets of purchased spam accounts, it’s likely that more sophisticated spammers and fake accounts go unidentified by our system"
The features collected are certainly outdated by now.
I'm sorry to be one of those people, but the low-contrast grey color used for the text is really annoying and makes me feel like I'm straining my eyes to read it! Why the fuck do designers encourage this? It is more readable as black, and I have a vague non-medically-informed sense that it might be better for people.
I am guessing that it makes sense that very high profile people will have a higher percentage of fake followers than an ordinary, small account would because many of these fake accounts are bots that are set up in order to get attention for whatever they're selling or promoting by interacting with high profile accounts.
So to get this straight - SparkToro have decided to burn their reputation down chasing a spurious irrelevant claim from someone else's merger. It'll never cease to amaze me how many people will take obvious trolling and treat is as reasonable. This is like conducting a serious analysis into whether the US election was stolen.
I hope they like law suits:
> Our analysis found that 19.42%, nearly four times Twitter’s Q4 2021 estimate, fit a conservative definition of fake or spam accounts
Ok great, I hope you have fun proving that in court. Especially the part where you have to prove that Twitter's definitions which you don't know match yours.
>SparkToro is a tiny team of just three
Then I applaud your bold decision to interfere with a $45Bn merger.
>Our definition (which may differ from Twitter’s own
Any lawyers in the house? How obviously do you have to renege on your libelous claims before you're in the clear?
I hope they like law suits:
> Our analysis found that 19.42%, nearly four times Twitter’s Q4 2021 estimate, fit a conservative definition of fake or spam accounts
Ok great, I hope you have fun proving that in court. Especially the part where you have to prove that Twitter's definitions which you don't know match yours.
>SparkToro is a tiny team of just three
Then I applaud your bold decision to interfere with a $45Bn merger.
>Our definition (which may differ from Twitter’s own
Any lawyers in the house? How obviously do you have to renege on your libelous claims before you're in the clear?
Main thing missing in this thread is that Bots can do things that doesn't involve actively tweeting. Things like liking, following, etc. Even clicking a hashtag has some effect on the Twitter algorithm. Furthermore, the most important metric for Twitter advertisers is Impressions. You primarily pay with impressions on Twitter as an advertiser. Does Twitter show ads to bots/fake/spam accounts if they match the #hashtag criteria or target audience to pump up the impression numbers?
Are these are the fake accounts that are lurking around without posting a tweet, but impacting the mDAU?
"Active" definition needs to be more precise than simply tweeting.
If I were an investor or an advertiser, I would drill into these details.
Are these are the fake accounts that are lurking around without posting a tweet, but impacting the mDAU?
"Active" definition needs to be more precise than simply tweeting.
If I were an investor or an advertiser, I would drill into these details.
So ~20% of Twitter accounts are fake; however, the last time I tried to create a Twitter account for legitimate purposes (asking some customer questions) I couldn't, just because I didn't want Twitter to have my phone number.
Fake is kind of a weird way to describe valid combinations of usernames and passwords that can log into Twitter.com. I would assume that if your account hasn't been disabled for whatever reason, it's a real Twitter account.
Twitter, however, does call these accounts "fake" and has further rules/policies concerning Misleading & Deceptive Identities.
https://help.twitter.com/en/rules-and-policies/twitter-imper...
Twitter, however, does call these accounts "fake" and has further rules/policies concerning Misleading & Deceptive Identities.
https://help.twitter.com/en/rules-and-policies/twitter-imper...
It's funny since I've had folks claim I must be a bot because I follow more folks than I have following me. It's just weird how these folks create metrics without much of a decent explanation. The fact of the matter is that bots are a problem that I think can only really be managed but not eliminated. The first thing is to make it botting not something that should be punished from the start but something that is used for common automated purposes just like how Twitter does it now with some bots being tagged as such for legitimate reasons but not blocked or shadowbanned.
I think it's instructive to note the difference between 'fake' accounts and 'spam'/'bot' accounts.
'Fake' accounts exist to fraudulently increase follower count. This is what the study is claiming to measure. These typically have low activity and engagement profiles.
'Spam' or 'bot' accounts, on the other hand, generally have high activity. Whether trying to influence political opinions, or engaging in astroturfing or phishing activities. They probably have very, very high ratios of replies to original tweets, and overall tweets to # followers.
'Fake' accounts exist to fraudulently increase follower count. This is what the study is claiming to measure. These typically have low activity and engagement profiles.
'Spam' or 'bot' accounts, on the other hand, generally have high activity. Whether trying to influence political opinions, or engaging in astroturfing or phishing activities. They probably have very, very high ratios of replies to original tweets, and overall tweets to # followers.
The methodology is not perfect but I think it raises a serious concern. I'm curious what other methodologies for identifying fake/spam accounts might be used. What about likes and retweets?
Perhaps Mr. Musks' team did an analysis of their own and saw a high number of potential fakes/bots, and that made them question the authenticity of Twitter's numbers.
If it can be shown with some accuracy that Twitter underreported the number of fake/spam accounts, how does this effect Musk's acquisition? Could he lower the price by saying you gave me incorrect data?
Perhaps Mr. Musks' team did an analysis of their own and saw a high number of potential fakes/bots, and that made them question the authenticity of Twitter's numbers.
If it can be shown with some accuracy that Twitter underreported the number of fake/spam accounts, how does this effect Musk's acquisition? Could he lower the price by saying you gave me incorrect data?
The biggest issue with any "study" like this is the definitions. What is a "bot"? What counts as "inactive"?
What about "bots" that aren't "bad bots"? For example, a news organization twitter account that automatically tweets out new articles. This is clearly a "bot", but is it a "fake account"?
Yes, most studies published specify what their own interpretation of these are, those definitions tend to be wildly inconsistent, and make any sort of comparison of the rest of the methodology impossible.
What about "bots" that aren't "bad bots"? For example, a news organization twitter account that automatically tweets out new articles. This is clearly a "bot", but is it a "fake account"?
Yes, most studies published specify what their own interpretation of these are, those definitions tend to be wildly inconsistent, and make any sort of comparison of the rest of the methodology impossible.
Elon already knew this of course. Buying Twitter was just an excuse to sell Tesla stock at a peak while avoiding suspicion. Now he’s teeing up his excuse to pull out. /conspiracy
The indicators of being a spambot they have in their post seem VERY iffy to me. "Not tweeting in the past 120 days", "Location set to a non resolving location", "Small number of followers", "default profile image", "No URL in bio or non-resolving URL in bio", "Not on many lists", "tweets in a different language than the person they're following" - Those all seem like extremely weak signals to me. My profile matches 6 of those, and I'm a human. I would like to see them hand-verify a subset of their results and see if their algorithm matches reality.
Also note that they define "active" differently than Twitter. They define "active" as having tweeted recently. Twitter gives spambot numbers as a percent of monetizable daily active users. I wonder if Twitter's given bot numbers are low because bots don't typically lurk or load ads. I can believe that the total bot count as a percentage of users or as a percentage of recently-tweeting-users is higher than 5%, but that only 5% of daily visitors seeing ads are bots.