That's why we checked whether advertisers are getting preferential filtering treatment. (They aren't according to our data.) Read my previous comment, or better yet read the paper where we also explain possible pitfalls in the analysis.
While I read such anecdotes, I cannot rely much on them for research. Instead, what we did for our paper was to collect thousands of filtered reviews, and check their characteristics against reviews Yelp publishes. We find that filtered reviews are shorter, more likely to be from one-time reviewers, more likely to be extreme (1- or 5-stars), and so on. I realize this does not constitute absolute proof but at least it's evidence based on data. Check the paper for more stats.
Beyond that, and as a counter-anecdote, I have read hundreds (if not thousands!) of filtered reviews over the past year and in my experience they are more likely to be fake. Furthermore, one implication of challenging our assumption is that you are suggesting consumers would be better off reading, and basing their decisions on filtered reviews since -- by your assumption -- they are more likely to be genuine. As I said, I cannot offer proof, but given what I have seen I believe the filter to be more likely to catch a fake review than a real review.
If there exists a material connection between the owner and the reviewer which is not disclosed in the review then the FTC considers this to be deceptive advertising. A material connection is defined to be something "that is, important to a consumer's decision to buy or use the product." See here: http://www.business.ftc.gov/documents/bus35-advertising-faqs...
So while you might not consider such a review as fake, my understanding is that the FTC would.
While I agree with you about the title of the article, it is not true at all that "all non-removed reviews are assumed, in the data, to be real." In fact, in our analysis we explicitly allow for the filter misclassifying reviews (that is to say filtering real reviews, or publishing fakes.)
(Mostly answering to the deleted grandparent here.)
"In the meantime, they benefit from charging businesses to allow positive reviews to appear and to suppress negative reviews (if they're doing that, it would explain why they don't focus on credibility)."
We try to address this issue in our paper. Take a look at Section 3.4. Short summary: we do not find any noticeable differences in filtering between advertisers and non-advertisers, but there are limitations to our analysis.
I don't know what goes into Yelp's filter either. I also agree that it can make mistakes. At the same time, I do think that the filter is more likely to catch a fake review than it is to catch a real review. This is the assumption our analysis relies on. (Not sure what I should be rethinking...)
Hi everyone, co-author of the paper here. I wanted to clarify the 20% figure. We actually have no way of telling how many, or which reviews are fake. We do not directly observe review fraud, and we clearly spell this out in the paper. The 20% figure represents the percentage of filtered reviews on Yelp. These are reviews that Yelp finds suspicious enough to not publish. Some filtered reviews may be fake, and some might just be false positives. Similarly, Yelp's filter might miss some fake reviews, and end up publishing them. (See http://www.yelp.com/faq#filter_wrong). I hope this makes the distinction between fake and filtered clear.
Our main goal with this paper was to analyze the economic incentives behind review fraud, and for this we used filtered reviews as a proxy for fake reviews. bobf provides a good summary of our key findings so I won't repeat them. For those interested, you can read more here: http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2293164
Are you saying it's Yelp's responsibility to protect unsophisticated users from themselves? Also, I am not quite sure "a lot of people" will behave as you describe.
True, Yelp got the filtering completely backwards in this instance. But, do you base any economic decision on two 5-star reviews from two Yelpers with 0 friends and 4 reviews combined?
You're right, statistical significance is a valid concern but error bars is not the way to go. Check the paper for a rigorous analysis. We are not drawing our conclusions from a single plot.