Almost all searches on my independent search engine are now from SEO spam bots(blog.searchmysite.net)
blog.searchmysite.net
Almost all searches on my independent search engine are now from SEO spam bots
https://blog.searchmysite.net/posts/almost-all-searches-on-my-independent-search-engine-are-now-from-seo-spam-bots/
377 コメント
Since everyone in this thread wants to jump down OP's throat about the quality of his web site, another interesting search engine is millionshort.com, which allows you to filter out the top N web sites from the results of your search. It's a great tool for looking past sites with good SEO; all you have to do is fiddle with the value of N.
For example, searching for "electronic music box" as /u/ajnin suggested, with the top 100K web sites removed from the results, filters out the following:
> These 23 sites were removed from your results:
> alibaba.com (1 result removed)
> aliexpress.com (1 result removed)
> allaboutcircuits.com (1 result removed)
> amazon.com (2 result removed)
> apple.com (1 result removed)
> bestreviews.com (1 result removed)
> ebay.com (1 result removed)
> etsy.com (2 result removed)
> facebook.com (1 result removed)
> instructables.com (2 result removed)
> lightinthebox.com (2 result removed)
> lumberjocks.com (1 result removed)
> mapquest.com (1 result removed)
> reverb.com (1 result removed)
> twitter.com (1 result removed)
> wikipedia.org (1 result removed)
> yelp.com (1 result removed)
> youtube.com (2 result removed)
And the top result ends up being https://midiguy.com/.
For example, searching for "electronic music box" as /u/ajnin suggested, with the top 100K web sites removed from the results, filters out the following:
> These 23 sites were removed from your results:
> alibaba.com (1 result removed)
> aliexpress.com (1 result removed)
> allaboutcircuits.com (1 result removed)
> amazon.com (2 result removed)
> apple.com (1 result removed)
> bestreviews.com (1 result removed)
> ebay.com (1 result removed)
> etsy.com (2 result removed)
> facebook.com (1 result removed)
> instructables.com (2 result removed)
> lightinthebox.com (2 result removed)
> lumberjocks.com (1 result removed)
> mapquest.com (1 result removed)
> reverb.com (1 result removed)
> twitter.com (1 result removed)
> wikipedia.org (1 result removed)
> yelp.com (1 result removed)
> youtube.com (2 result removed)
And the top result ends up being https://midiguy.com/.
This made me curious to try that search engine so I typed "electronic music box" (first thing that came to mind). As far as I can tell none or the 10+ pages of results include all those 3 words. I mean, you might not have any relevant sites in your database (likely if there are only 1000 sites or so as another of your blog posts imply), and I understand you want to show some result to the user, but if I want irrelevant links I might as well go to google.com...
You mention the "Dead Internet Theory" (not heard that phrase before!).
I agree: the WWW Internet is dead, that is your problem. No-one visits websites anymore, everyone has moved to the 10 biggest websites and all data is now siloed there.
If I want to search for something topical and relevant, I go to Facebook, Twitter, Reddit, HackerNews, Instagram, Google Maps, Discord etc.
The general Internet is dead: it's just legacy content and spam.
If you think it's bad for you, imagine what it is like for Google Search! Their entire business is indexing a medium which no longer has any relevancy. People complain that Google no longer delivers good results. But what can Google do? The "good content" is no longer available for them to index.
Want to become rich? Make a search engine which indexes the fresh relevant data from the big siloed websites, and ignores the general dead Internet.
I agree: the WWW Internet is dead, that is your problem. No-one visits websites anymore, everyone has moved to the 10 biggest websites and all data is now siloed there.
If I want to search for something topical and relevant, I go to Facebook, Twitter, Reddit, HackerNews, Instagram, Google Maps, Discord etc.
The general Internet is dead: it's just legacy content and spam.
If you think it's bad for you, imagine what it is like for Google Search! Their entire business is indexing a medium which no longer has any relevancy. People complain that Google no longer delivers good results. But what can Google do? The "good content" is no longer available for them to index.
Want to become rich? Make a search engine which indexes the fresh relevant data from the big siloed websites, and ignores the general dead Internet.
Ona tangential note, I remember a time when Google had the option to search only for 'discussions'. The results were amazing and accurate as it scoured online forums. Almost all issue I had (was following the rooting scene closely back then) were quickly resolved. Then suddenly it got removed for reasons unknown to me. Anyone knows if it's replicatable today?
> I didn’t notice at first because the web analytics only shows real users, and the unusual activity could only be seen by looking at the server logs.
Sounds like everyone blocking analytics (Plausible in this case), e.g. myself just now, is lumped in with spam bots.
Of course, analytics blocking can’t meaningfully swing the ~99.99% statistic.
Sounds like everyone blocking analytics (Plausible in this case), e.g. myself just now, is lumped in with spam bots.
Of course, analytics blocking can’t meaningfully swing the ~99.99% statistic.
I'm disappointed that Search My Site isn't seeing many legitimate viewers.
Just wanted you to know that I'm a fan. I love reading peoples personal websites, and Search My Site has been great for discoverability. I visit the Newest Pages and Browse Sites pages once or twice a week to check out the new sites being indexed.
I don't know what the answer is to the spam bots, but you do have some real visitors out there. :)
Just wanted you to know that I'm a fan. I love reading peoples personal websites, and Search My Site has been great for discoverability. I visit the Newest Pages and Browse Sites pages once or twice a week to check out the new sites being indexed.
I don't know what the answer is to the spam bots, but you do have some real visitors out there. :)
This guy throws multiple reasons/conspiracies out there on why the website is really struggling to gain literally any sort of traction. Web is all bots, search engines not promoting competitors and being drowned out by SEO spam, yet he's failing to see the most obvious reason... the reason nearly all websites don't gain traction...
Because it's a bad website. It provides no value to the user. I put in a few search terms and had no relevant search results back. What use is a search engine that can't find what I'm searching for?
Maybe if that was improved he may see traction.
Because it's a bad website. It provides no value to the user. I put in a few search terms and had no relevant search results back. What use is a search engine that can't find what I'm searching for?
Maybe if that was improved he may see traction.
Search traffic has always been mostly automated spam bots.
Even back in the Open Directory Days when we powered part of search.netscape.com I estimated 80+% of all search traffic was automated. At least most of it self-identified with the same Java useragent.
Later when working Topix, despite being a news search engine, most traffic was bot traffic. Most included the word “mortgage” in the query. Topix specialized in localized content, and that was very popular for SEO scrapers.
Lastly at Blekko, I estimate 90+% of traffic was automated. By then maybe half or more learned to change the user agent. Most used HTTP/1.0, a dead giveaway as no browser still uses 1.0. This was a major aspect in Blekko's load shedding strategy. If the servers started to get overloaded, we'd start bouncing suspected bot traffic to a redirect that would show in the logs. If there was a human with a modern browser running javascript on the other end, would get redirect to a link that wouldn't get bounced. I would check the logs weekly to see if any humans got caught. None ever did. This was a huge monetary savings, you only need 1/10th the servers if you can safely ignore the bots.
Often it's endless repetition of the same keywords in a random order with a place name appended, or prepended, or inserted. over and over. Often variations on known monetizatable SEO keywords. However, much of it doesn't make any sense.
I don't have any insight into Google's numbers but I would conservatively estimate 95% or more of all their queries are automated bots and not humans. And the level of spy-vs-spy going on for Google CPU resources vs SEO bots is probably pretty evolved by now. I stopped tracking many years ago when Google switched to densely packed obfuscated javascript for page renders. Maybe this is part of why automated queries are so high across the web, maybe google is too hard to crack for most.
Even back in the Open Directory Days when we powered part of search.netscape.com I estimated 80+% of all search traffic was automated. At least most of it self-identified with the same Java useragent.
Later when working Topix, despite being a news search engine, most traffic was bot traffic. Most included the word “mortgage” in the query. Topix specialized in localized content, and that was very popular for SEO scrapers.
Lastly at Blekko, I estimate 90+% of traffic was automated. By then maybe half or more learned to change the user agent. Most used HTTP/1.0, a dead giveaway as no browser still uses 1.0. This was a major aspect in Blekko's load shedding strategy. If the servers started to get overloaded, we'd start bouncing suspected bot traffic to a redirect that would show in the logs. If there was a human with a modern browser running javascript on the other end, would get redirect to a link that wouldn't get bounced. I would check the logs weekly to see if any humans got caught. None ever did. This was a huge monetary savings, you only need 1/10th the servers if you can safely ignore the bots.
Often it's endless repetition of the same keywords in a random order with a place name appended, or prepended, or inserted. over and over. Often variations on known monetizatable SEO keywords. However, much of it doesn't make any sense.
I don't have any insight into Google's numbers but I would conservatively estimate 95% or more of all their queries are automated bots and not humans. And the level of spy-vs-spy going on for Google CPU resources vs SEO bots is probably pretty evolved by now. I stopped tracking many years ago when Google switched to densely packed obfuscated javascript for page renders. Maybe this is part of why automated queries are so high across the web, maybe google is too hard to crack for most.
One day we'll have an internet for humans exclusively. On another note, with 160K requests / day from bots you could of course simply block the bots structurally assuming they are nice enough to identify themselves. Block all of AWS and Google, Russia, China, NK and a couple of other bot hot spots and the service may well become more successful for regular users because they get faster results. Bots can afford to wait, humans are often impatient. And with 2 hits / second by bots that may well become a factor.
This is for comment spam.
It's trying to find a long tail of popular but not top listed blogs for the purpose of posting comments with the much desired links to the SEO target.
It's trying to find a long tail of popular but not top listed blogs for the purpose of posting comments with the much desired links to the SEO target.
If the internet is dead, is there anything left that's "alive"? The mobile app stores are also filled with crap[0] and it seems that the ratio of spam content vs real content is getting close to infinity.
[0]: https://youtu.be/E8Lhqri8tZk - 1,500 Slot Machines Walk into a Bar: Adventures in Quantity Over Quality
[0]: https://youtu.be/E8Lhqri8tZk - 1,500 Slot Machines Walk into a Bar: Adventures in Quantity Over Quality
I have had very interesting conversations with people who are "casual" users of internet. They are still finding the results of the likes of Google, bing and duckduckgo perfectly suitable. Maybe it's most of us here who have different needs to what's available.
Mojeek member here. We have always had a high level of spam bots; as any search engine/service will have. It's a constant battle to fend off new bots; folks can always use try out our API rather than freeloading, and some do. Many obviously do not. We are taking a look at whether things have also changed for us since mid-April 2022.
That's really interesting... and sad. For what it's worth, I've noticed comment bots dramatically increase over the last year too. They have always been there, but looking at Reddit, YouTube, etc, now there seem to be 10x more than there were a few years earlier. Even on HN it has gotten worse.
i recently built a habitat for spam bots, they eventually found it and now post peacefully
https://upstairs.treehouse.telnet.asia/pharm/cylohexapine
https://upstairs.treehouse.telnet.asia/pharm/cylohexapine
I had to put my search engine behind Cloudflare to deal with this. Like the volume grew to about 10x the traffic I saw sitting at the front page of Hacker News for a full week.
If this was the spam for a search engine (almost) nobody uses, it makes you wonder how much abuse the major search engines face
My key takeaways:
1. Almost all searches on my independent search engine are now from SEO spam bots
2. In summary, if they break through the current reverse proxy level protection, options include an invisible ReCAPTCHA (but given I’ve sometimes 160,000 requests a day I’d be well over the 1,000,000 a month free tier limit), requiring JavaScript as per the web analytics or some Cross Site Request Forgery style protection (but those would place much more load on the servers), or CloudFlare (but the searchmysite.net spider is still currently blocked by CloudFlare as per Some of the challenges of building an internet search)
3. If you were into conspiracy theories you could claim that the major search engines were trying to stifle the competition, but a more realistic explanation is simply that searchmysite.net is being drowned out by SEO spam
4. If I’d had a decent amount of real users visiting and never returning I could reasonably conclude that updating the blog wasn’t the most productive use of my time and effort, but without any real users in the first place it is hard to gauge whether people like it or not
My own independent search engine, https://www.locserendipity.com, is seeing similar trends.
1. Almost all searches on my independent search engine are now from SEO spam bots
2. In summary, if they break through the current reverse proxy level protection, options include an invisible ReCAPTCHA (but given I’ve sometimes 160,000 requests a day I’d be well over the 1,000,000 a month free tier limit), requiring JavaScript as per the web analytics or some Cross Site Request Forgery style protection (but those would place much more load on the servers), or CloudFlare (but the searchmysite.net spider is still currently blocked by CloudFlare as per Some of the challenges of building an internet search)
3. If you were into conspiracy theories you could claim that the major search engines were trying to stifle the competition, but a more realistic explanation is simply that searchmysite.net is being drowned out by SEO spam
4. If I’d had a decent amount of real users visiting and never returning I could reasonably conclude that updating the blog wasn’t the most productive use of my time and effort, but without any real users in the first place it is hard to gauge whether people like it or not
My own independent search engine, https://www.locserendipity.com, is seeing similar trends.
Well, the first two links loaded for a search for "magic the gathering" are 404s. The "Random" link at the bottom 403s. The search engine feels broken.
I run a data aggregation company that has a fairly advanced scraping infrastructure for collecting data across the web. Having built the scraping side, I'm pretty familiar with most of the strategies for avoiding bot detection.
Coming from that perspective, detecting and stopping at least the majority of bots out there is fairly doable, and I put together a rudimentary thing for a side project.
The core of it uses an IP API for looking up the requesting IP to identify the country and if it's coming from a data center, VPN, Tor, etc. If it passes that, I trigger Google Captcha to show up. Lastly, I track IPs that make it through and have some basic rules in place to try to detect patterns and block offenders that way.
There's a bunch more stuff you can check for, but the core of it is basically filtering out data center traffic to minimize the requests going to Google Captcha.
Coming from that perspective, detecting and stopping at least the majority of bots out there is fairly doable, and I put together a rudimentary thing for a side project.
The core of it uses an IP API for looking up the requesting IP to identify the country and if it's coming from a data center, VPN, Tor, etc. If it passes that, I trigger Google Captcha to show up. Lastly, I track IPs that make it through and have some basic rules in place to try to detect patterns and block offenders that way.
There's a bunch more stuff you can check for, but the core of it is basically filtering out data center traffic to minimize the requests going to Google Captcha.
Complete SEO noob here. Can someone help explain what these bots are trying to achieve? There is mention in the blog that they're trying to uncover ad free content.
IMHO what you should try is excluding all sites with excessive third-party cookies, sluggish performance, and too many ads. That will slice the index down by 80% probably but it would be a really nice thing to see. It might push out low quality SEO results for a couple of years.
In late April up to now, Wiby (a small mostly unheard of search engine) began having the exact same issue. Tens of thousands of the exact same type of "powered by..." requests coming from thousands of IPs. They are using a tool called QHub.
Spammers badly need spam-free content so they can mix some legitimate links with the junk they spew.
One great Black Hat SEO trick is to find where your competitors are getting clean links and insert your own links there so they do your spamming for you.
One great Black Hat SEO trick is to find where your competitors are getting clean links and insert your own links there so they do your spamming for you.
Random tangent related to SEO
I am so annoyed with random overseas companies faking local businesses by SEO
They will get your call and you can tell where you're talking to immediately. Then they get a quote that's increased to get their cut... That's then carried out by an actual local company. It's annoying because the websites at the top of the search appear local.
A specific example is when you're looking for a towing company.
Type in your state/city towing company, more than likely the top results/websites pinned to GMaps are not local-based.
They will get your call and you can tell where you're talking to immediately. Then they get a quote that's increased to get their cut... That's then carried out by an actual local company. It's annoying because the websites at the top of the search appear local.
A specific example is when you're looking for a towing company.
Type in your state/city towing company, more than likely the top results/websites pinned to GMaps are not local-based.
I think many people in the comments here, and most users, are missing that you index a SMALL subset of the web. This leads to people running a default test search, finding no results, and concluding your search engine is bad, and leaving.
While you imply that in the search page, obviously it's not clear enough.
Maybe add "this search engine only searches a small set of user submitted sites. Click <here> for the list. Or <here> to add your site."
While you imply that in the search page, obviously it's not clear enough.
Maybe add "this search engine only searches a small set of user submitted sites. Click <here> for the list. Or <here> to add your site."
>I noted that there had been multiple weeks where not one single real person had visited a single blog entry for the whole week
The site is not on https://searchengine.party/ nor on seirdy.one's overview. Apart from the blog, how could users find that engine?
Is there some place where new search engines are announced and where new search engines band together to make themselves heard?
The site is not on https://searchengine.party/ nor on seirdy.one's overview. Apart from the blog, how could users find that engine?
Is there some place where new search engines are announced and where new search engines band together to make themselves heard?
I created a temporary email service that was being used by about 10k users / week. Then several weeks ago, the number of users started growing like crazy up to about 60k users a day. Then we checked the recent email activity and 60k / 65k emails were from a social networking site.
Seems our service was being used to create fake bot accounts. The newly created accounts were obvious fakes. Rather than deal with the issue, we just shut the service off.
Seems our service was being used to create fake bot accounts. The newly created accounts were obvious fakes. Rather than deal with the issue, we just shut the service off.
This is an awsome website that I was not aware of!
However, they are also systematically feeding you their footprint lists. I imagine you could put together a footprint blacklist pretty quickly, and just stop returning results for any obvious spam queries like those containing "powered by wordpress".
It's not a very elegant solution I'll admit. It won't stop the bots from trying, and you may have to circle back periodically to add new footprints as they surface. But it's a potentially quick and easy way to stop rewarding their efforts, and the blackhat world is pretty used to burning out their resources so hopefully they will figure out it's a dead end and move on.