This is not how to crawl webpages. He started with the Alexa list. Those are not necessarily domain names of servers serving webpages. I would guess that some of the request to cease crawling came from some of these listings. Working from the Alexa list he would have been crawling some of the darkest underbelly of the web: ad servers and bulk email services.
His question: "Who gets to crawl the web?" is an interesting one though.
Do not assume that Googlebot is a smart crawler. Or smarter than all others. The author of Linkers and Loaders posted recently on CircleID about how dumb Googlebot can be.
There is no such thing as a smart crawler. All crawlers are stupid. Googlebot resorts to brute force more often than not.
Theoretically no one should have to crawl the web. The information should be organised when it is entered into the index.
Do you have to "crawl" the Yellow Pages? Are listings arranged by an "algorithm"? PageRank? 80/20 rules?
Nothing wrong with those metrics; except of course that they can be gamed trivially, as experiments with Google Scholar have shown. But building a business around this type of ranking? C'mon.
If the telephone directories abandoned alpha and subject organisation for "popularity" as a means of organisation it would be total chaos. Which is why "organising the world's information" is an amusing mission statement when your entire business is built around enabling continued chaos and promoting competition for ranking.
Even worse are companies like Yelp. It's blackmail.
If the information was organised, e.g., alphabetically and regionally, it would be a lot easier to find stuff. Instead, search engines need to spy on users to figure out what they should be letting users choose for themselves. Where "user interfaces" are concerned, it is a fine line between "intuitive" and "manipulative".
The people who run search engines and directory sites are not objective. They can be bought. They want to be bought.
This brings quality down. As it always has for traditional media as well. But it's much worse with search engines.
His question: "Who gets to crawl the web?" is an interesting one though.
Do not assume that Googlebot is a smart crawler. Or smarter than all others. The author of Linkers and Loaders posted recently on CircleID about how dumb Googlebot can be.
There is no such thing as a smart crawler. All crawlers are stupid. Googlebot resorts to brute force more often than not.
Theoretically no one should have to crawl the web. The information should be organised when it is entered into the index.
Do you have to "crawl" the Yellow Pages? Are listings arranged by an "algorithm"? PageRank? 80/20 rules?
Nothing wrong with those metrics; except of course that they can be gamed trivially, as experiments with Google Scholar have shown. But building a business around this type of ranking? C'mon.
If the telephone directories abandoned alpha and subject organisation for "popularity" as a means of organisation it would be total chaos. Which is why "organising the world's information" is an amusing mission statement when your entire business is built around enabling continued chaos and promoting competition for ranking.
Even worse are companies like Yelp. It's blackmail.
If the information was organised, e.g., alphabetically and regionally, it would be a lot easier to find stuff. Instead, search engines need to spy on users to figure out what they should be letting users choose for themselves. Where "user interfaces" are concerned, it is a fine line between "intuitive" and "manipulative".
The people who run search engines and directory sites are not objective. They can be bought. They want to be bought.
This brings quality down. As it always has for traditional media as well. But it's much worse with search engines.