Google: We knew the web was big...(googleblog.blogspot.com)
googleblog.blogspot.com
Google: We knew the web was big...
http://googleblog.blogspot.com/2008/07/we-knew-web-was-big.html
25 comments
"Strictly speaking, the number of pages out there is infinite -- for example, web calendars may have a "next day" link, and we could follow that link forever, each time finding a "new" page."
Strictly speaking, even that would keep the number of pages finite, even if very, very big. Any amount of data you can produce with any number of finite computers (with finite and bounded memory -there are practical bounds for memory sizes), will always be finite.
Strictly speaking, even that would keep the number of pages finite, even if very, very big. Any amount of data you can produce with any number of finite computers (with finite and bounded memory -there are practical bounds for memory sizes), will always be finite.
That's the same as saying an infinite data structure is finite; just because it cannot be practically evaluated before the heat death of the universe doesn't mean it isn't theoretically infinite.
No: with a finite amount of memory, the amount of information you can generate in infinite time is finite.
You're assuming the systems themselves wouldn't grow or be upgraded over time.
No; I'm saying the amount of information at any given time is finite, which is what the original article was talking about. But even if there is growth in memory size (say that memory size increases linearly), the information you get at any specific instant is finite (even if unbounded)... just as you have infinite natural numbers but none of them has an infinite value. But anyway, with finite space even in infinite time you couldn't build an infinite memory.
And also... the original article was talking about the size of the system at a specific instant, so growth doesn't matter here.
And also... the original article was talking about the size of the system at a specific instant, so growth doesn't matter here.
Given infinite time, the number of pages that could be generated is infinite.
With a finite amount of memory, the amount of information you can generate in infinite time is finite.
You don't have to store every page you generate.
No, but at least I suppose that pages are bounded by size, because of the memory of the generator, or of the viewer. Say the maximum size of a page you can build is N, then there is a maximum finite value of different pages / urls you can fit in N too. So, even if you have infinite time, with your finite computers you'll have finite amounts of pages.
Now, if you are going to have pages and urls that are bigger in size than the memory of even the biggest computers in the world, so that you can't hold them in memory at any time, then the point above isn't useful for you (and in this case you are right). But urls bigger than memory are probably useless, even if webpages bigger than memory aren't.
And if you have to have urls that are finite... then my point holds: even in infinite time, you won't have infinite amounts of urls.
Now, if you are going to have pages and urls that are bigger in size than the memory of even the biggest computers in the world, so that you can't hold them in memory at any time, then the point above isn't useful for you (and in this case you are right). But urls bigger than memory are probably useless, even if webpages bigger than memory aren't.
And if you have to have urls that are finite... then my point holds: even in infinite time, you won't have infinite amounts of urls.
You don't need to keep in memory the page you generate either.
Well if you don't keep in memory the page nor the url you generate ever (either in the publisher or in the viewer), can it be used in any meaningful way? But if you don't keep urls in memory you could possibly generate as many urls as you wish in infinite time...
Only if they take no energy to generate.
TechCrunch is being cryptic (http://www.techcrunch.com/2008/07/25/googles-misleading-blog...)
“Google also says 'But we’re proud to have the most comprehensive index of any search engine.'
That may be true today, but it probably won’t be true next week (check back here then). Google knows that as well as we do, and that’s why they posted this today."
So, is it Yahoo or Live? If it's one of them, why would Google know it so well?. Any thoughts on what this could be? One of the TC commenters thinks its MSFT indexing facebook ...
“Google also says 'But we’re proud to have the most comprehensive index of any search engine.'
That may be true today, but it probably won’t be true next week (check back here then). Google knows that as well as we do, and that’s why they posted this today."
So, is it Yahoo or Live? If it's one of them, why would Google know it so well?. Any thoughts on what this could be? One of the TC commenters thinks its MSFT indexing facebook ...
TechCrunch is being cryptic
Yeah, this is surprising... especially because Arrington would never lie to drive traffic to his site. Oh wait. TechCrunch.
Yeah, this is surprising... especially because Arrington would never lie to drive traffic to his site. Oh wait. TechCrunch.
well allow me to modify my robots.txt so you can be at infinity-1. since you pricks decided to hijack your role as impartial algorithmic search to walled-garden with knol, i want nothing to do with you.
every major website that handles referrals seems to bite the poison fruit of capturing traffic...now google has to. oh well, maybe clusty search will have to do
every major website that handles referrals seems to bite the poison fruit of capturing traffic...now google has to. oh well, maybe clusty search will have to do
Great! So now, all you have to do is give some sort of a comparison that lets people visualise exactly how big an area 50,000 times the size of U.S. is. May I suggest a metric based on whales?