Does Adding Many Tags to an Instagram Photo Maximize the Number of Likes?(minimaxir.com)
minimaxir.com
Does Adding Many Tags to an Instagram Photo Maximize the Number of Likes?
http://minimaxir.com/2014/03/hashtag-tag/
5 comments
But where did they get more traffic from to begin with? You're implying they got traffic from "fill-all-the-keywords strategies", which is more-or-less what the analysis is advocating. That may not be true, but your alternative hypothesis isn't actually an alternative. At least not as I understand it.
That said, I do think controlling for follower counts would be informative. It would help tell you if there's actually something else at play (e.g. once you control for follower count, there's no more correlation). I suspect it could also tell you the # of hashtags is even more important than currently shown because it also helps build your base of followers, by reaching more new visitors.
E.g. If increased hashtags increase my number of favorites, they also increase my number of followers by some percentage. Then each successive post has a larger direct reach of followers, in addition to the new visitor reach from hashtag discovery. A single hashtag could actually have a compounding value.
That said, I do think controlling for follower counts would be informative. It would help tell you if there's actually something else at play (e.g. once you control for follower count, there's no more correlation). I suspect it could also tell you the # of hashtags is even more important than currently shown because it also helps build your base of followers, by reaching more new visitors.
E.g. If increased hashtags increase my number of favorites, they also increase my number of followers by some percentage. Then each successive post has a larger direct reach of followers, in addition to the new visitor reach from hashtag discovery. A single hashtag could actually have a compounding value.
My alternative theory might explain it: Bad actors who stuff their photos with 30 keywords also use shill accounts to like them.
My comment was not very clear. I’ll try again.
For simplicity let’s suppose that there are only two types of users:
* Some users have an account in every social network, have a very popular blog, perhaps pay for some advertisement, fill in all the keywords / tags fields, and use all the SEO they can.
* Some users take the photo, and just copy the nouns of the title to the tags field.
If this model is correct, most of the photos with lots of tags are from users with lots of followers and that is the cause of the additional likes.
The original (implicit) idea of the article is that the photos with more tags appear in more search results, and this is the cause of the additional likes. [I hope this simplified version is an accurate enough.]
Real life is usually more complicated, there are more than two types of users, and perhaps both effects are present. I like the analysis of the data, but I’m not sure if this can be transformed directly in a strategy to gain followers. Perhaps this is a consequence of cargo cult SEO.
How can we setup an experiment to distinguish between the two effects?
For simplicity let’s suppose that there are only two types of users:
* Some users have an account in every social network, have a very popular blog, perhaps pay for some advertisement, fill in all the keywords / tags fields, and use all the SEO they can.
* Some users take the photo, and just copy the nouns of the title to the tags field.
If this model is correct, most of the photos with lots of tags are from users with lots of followers and that is the cause of the additional likes.
The original (implicit) idea of the article is that the photos with more tags appear in more search results, and this is the cause of the additional likes. [I hope this simplified version is an accurate enough.]
Real life is usually more complicated, there are more than two types of users, and perhaps both effects are present. I like the analysis of the data, but I’m not sure if this can be transformed directly in a strategy to gain followers. Perhaps this is a consequence of cargo cult SEO.
How can we setup an experiment to distinguish between the two effects?
Just keep in mind an R^2 of 15 isn't that high, so these variables are only accounting for a small portion of the variance in the model.
I think there's an important control here: number of followers per user. It'd be interesting to see how followers would affect the model, especially because you could assume that "knowing to add more hashtags" likely correlates with "use of the platform," which they likely correlates with "increased # followers," who can then produce more likes on each photo.
EDIT: Another thought: if you're looking at # likes, you might want to go for a negative binomial model instead of an OLS regression, because it'll account for the dependent variable being a "count" measure.
I think there's an important control here: number of followers per user. It'd be interesting to see how followers would affect the model, especially because you could assume that "knowing to add more hashtags" likely correlates with "use of the platform," which they likely correlates with "increased # followers," who can then produce more likes on each photo.
EDIT: Another thought: if you're looking at # likes, you might want to go for a negative binomial model instead of an OLS regression, because it'll account for the dependent variable being a "count" measure.
Well, what is your expected outcome? If you're just measuring to see what happens, you're not really going to know unless you do a larger experiment. If you get more specific it's easier to tell what exactly you need to do in order to get an accurate result.
If somebody repeats that cliche in person to me any time soon, I will stick a fork in their head. And I thought the "curse of dimensionality" and "embarassingly parellizable" were annoying.
edit: Although, yes, certainly a valid point.
edit: Although, yes, certainly a valid point.
That was a lot of analysis that seems pretty useless to me. What action would you take based on that information? I know that my photos get more likes if I use hashtags. If I don't hashtag a photo I only get likes from my existing followers. If I do put on hashtags I get some likes from people that don't follow me. I'm sure they are searching for photos by hashtag. I myself find photos and people to follow based on hashtag searches. The bulk of likes I receive are always from my followers and hashtags make no difference to them. But hashtags bring new followers.
(This post assumes that one is trying to maximize likes. There are obviously different, and arguably superior, ways of using a social network)
The value of analysis like this is proving, codifying, and quantifying things that are intuitively believed to be true. Or, alternatively, disproving those things.
One might believe that "more hashtags" causes "more likes", and even have a model in which "more hashtags" causes "more likes", but that doesn't mean that it's actually true.
Before doing the analysis, the result could have come out either way. The author could have found that "more hashtags" did not cause (or at least, was not correlated with) "more likes". In this case, one could have taken the action of not filling in 30 hashtags, saving effort.
And, for that matter, there's a lot of variance in average likes that's not accounted for by hashtag count. Further analysis might actually prove that some other factor (such as age of poster) influences both like count and hashtag count, and accounting for that other factor, hashtags might actually be useless or detrimental.
-----
As a general philosophy, intuitively knowing "something is true" is not nearly as good as being able to point at data, and show general evidence that the thing is actually likely to be true.
The value of analysis like this is proving, codifying, and quantifying things that are intuitively believed to be true. Or, alternatively, disproving those things.
One might believe that "more hashtags" causes "more likes", and even have a model in which "more hashtags" causes "more likes", but that doesn't mean that it's actually true.
Before doing the analysis, the result could have come out either way. The author could have found that "more hashtags" did not cause (or at least, was not correlated with) "more likes". In this case, one could have taken the action of not filling in 30 hashtags, saving effort.
And, for that matter, there's a lot of variance in average likes that's not accounted for by hashtag count. Further analysis might actually prove that some other factor (such as age of poster) influences both like count and hashtag count, and accounting for that other factor, hashtags might actually be useless or detrimental.
-----
As a general philosophy, intuitively knowing "something is true" is not nearly as good as being able to point at data, and show general evidence that the thing is actually likely to be true.
Interesting approach. I wonder what the result would look like if time were taken into account - photos that are around longer will have a higher probability of more likes by default. It would be easy to correct for this by only considering photos that have been posted for the same length of time.
What is the distribution of likes over time on Instagram anyway?
What is the distribution of likes over time on Instagram anyway?
Time is definitely a factor. Most of the Likes are made within 24 hours, but then they taper off. I'll look more into that for a future blog post.
(Out of curiosity, I did rerun all the charts in the post with the most-recent photos removed. The resulting charts were relatively unchanged.)
(Out of curiosity, I did rerun all the charts in the post with the most-recent photos removed. The resulting charts were relatively unchanged.)
Brace yourselves, even more hashtags are coming.
Am I the only one who wishes the time and effort spent into making this post was used towards something more important/meaningful?
It's actually not that time consuming to make the charts once you've got the templating down. :)
[deleted]
Alternative theory (without any data :)):
Perhaps most of the photos with 30 tags are dew to semiautomatic fill-all-the-keywords strategies in SEO attempts. So most of them have naturally more exposure than photo from individuals that only use a few tags (like “dog” and “sleeping”).
Then they get more likes because they have more traffic, not because they have more tags
Proposed experiment:
* Pick 70 new photos, and separate them in 7 groups of 10 photos at random.
* Post them with 0; 5; 10; ... ; 30 tags and wait ...
* Count the likes for each group
* Kill outliers? 10 is too small? Find someone with more knowledge of statistics to check the experiment.