Web 3.0: The Semantic Web, a short film(kateray.net)
kateray.net
Web 3.0: The Semantic Web, a short film
http://kateray.net/film
3 comments
I just don't think people will be willing to add the semantic data. Instead of the semantic web, I think we will see more and more services that can extract information automatically.
For example, if you tweet or blog about a place, it seems doable to automatically realize that it is a place. You don't have to tweet "Going to <semantic:place>Joe's Bar</semantic:place>". From knowing your default location, a clever computer could recognise "Joe's Bar" as a place, because your city has a place called "Joe's Bar".
I think there are already examples of this cropping up here and there (and Yahoo has a service for places you can use).
For example, if you tweet or blog about a place, it seems doable to automatically realize that it is a place. You don't have to tweet "Going to <semantic:place>Joe's Bar</semantic:place>". From knowing your default location, a clever computer could recognise "Joe's Bar" as a place, because your city has a place called "Joe's Bar".
I think there are already examples of this cropping up here and there (and Yahoo has a service for places you can use).
I couldn't disagree more. I see your viewpoint, but I think think that's extrapolating based on our current perspective -- initiatives like LinkedData, which don't work.
If you give people a tool that can do it clearly, flexibly and easily, they will blow your mind.
If you give people a tool that can do it clearly, flexibly and easily, they will blow your mind.
Tichy is correct. The semantic web is essentially a big knowledge management effort that's doomed to failure (like most KM efforts) because xyz Joe Random web page maker really doesn't see any incentive to go through and spend a couple hours of his time marking his page about futons up for the SW. It's a waste of his time.
Even if he did, by some miracle do that. There's no way somebody else can make much more use of that information beyond what they already can through other, lower effort facilities like search for example. It requires powerful, web-scale reasoning engines to even make sense of that, none of which we have today.
Supposing even 5 years from now, some startup is founded to do the reasoning, say Cuil 2.0, what do I get for it? Better page retrieval? Is it 2% better? Is it 30% better?
It better be 30% better, because 2% better isn't worth hundreds of billions of aggregate dollars for everybody to go through and mark up all their pages to make them SW ready. And since nobody really knows (there isn't a useful reasoning engine that does anything that you can't do with search in the web space that I'm aware of, and most of the ones I've seen provide results far worse than search, Wolfram Alpha is probably the closest to a half-working semantic reasoning system you can find, and it's not like it's taken the world by storm), nobody will even bother -- a classic catch-22.
The SW is just a repackaging of old broken AI ideas. We already know systems built on essentially the same principles don't do much for the average person, the SW is just a doubling down on the same bad ideas -- "it didn't work before because it wasn't big enough!"
Even if he did, by some miracle do that. There's no way somebody else can make much more use of that information beyond what they already can through other, lower effort facilities like search for example. It requires powerful, web-scale reasoning engines to even make sense of that, none of which we have today.
Supposing even 5 years from now, some startup is founded to do the reasoning, say Cuil 2.0, what do I get for it? Better page retrieval? Is it 2% better? Is it 30% better?
It better be 30% better, because 2% better isn't worth hundreds of billions of aggregate dollars for everybody to go through and mark up all their pages to make them SW ready. And since nobody really knows (there isn't a useful reasoning engine that does anything that you can't do with search in the web space that I'm aware of, and most of the ones I've seen provide results far worse than search, Wolfram Alpha is probably the closest to a half-working semantic reasoning system you can find, and it's not like it's taken the world by storm), nobody will even bother -- a classic catch-22.
The SW is just a repackaging of old broken AI ideas. We already know systems built on essentially the same principles don't do much for the average person, the SW is just a doubling down on the same bad ideas -- "it didn't work before because it wasn't big enough!"
Better page retrieval? This is what a lot of people just don't understand. It's not "what exists plus 30%". It's a different equation entirely.
I live in the Czech Republic. A young girl recently put a note on our train station board offering babysitting services. Why didn't she use the web? Because she can't. Until someone makes a Czech babysitting database, she can't. Because no one would ever find it. Text doesn't work that way.
How many web sites are: dating, real estate, jobs, etc? That's data. So why don't they put it on the web as text? Because no one would ever find it. Text doesn't work that way.
Whoever fixes this, changes it all. They take out Google. They take out Monster.com. They take out Match.com. They take out Facebook.
Whoever can do this will change everything.
We've spent 15 years building the mountain of the web on the wrong construct. We need to rebuild it.
You think it's doomed to failure but I'm telling you concretely, you'll see it within a year from today.
I live in the Czech Republic. A young girl recently put a note on our train station board offering babysitting services. Why didn't she use the web? Because she can't. Until someone makes a Czech babysitting database, she can't. Because no one would ever find it. Text doesn't work that way.
How many web sites are: dating, real estate, jobs, etc? That's data. So why don't they put it on the web as text? Because no one would ever find it. Text doesn't work that way.
Whoever fixes this, changes it all. They take out Google. They take out Monster.com. They take out Match.com. They take out Facebook.
Whoever can do this will change everything.
We've spent 15 years building the mountain of the web on the wrong construct. We need to rebuild it.
You think it's doomed to failure but I'm telling you concretely, you'll see it within a year from today.
Of course if she put an ad for "babysitting" on the web, people would find it. There is nothing "semantic" about it. You simply search for the word "babysitting".
More likely, she doesn't have a computer or the parents don't have computers.
More likely, she doesn't have a computer or the parents don't have computers.
Or, (and I say this living in the Czech Rep, and I just did a search on this country's top search site which is not google) if she put up a web page, nobody would find it, unless they wade through 30 pages of results. Unless she is a SEO wizard, it actually makes more sense for her to target locally (train station in the town she lives in) in a way that people will see it. Or, it could be financially motivated. (Hosting costs, etc...)
Have you tried searching on Google Maps? I think they already automatically extract addresses from Web sites. So if you make a babysitting web site with your address, it should work.
Also not sure how semantic web would help with targeting locally.
Also not sure how semantic web would help with targeting locally.
I haven't tried google maps. But, neither would any locals out here. Google is number 3 out here for searches.
And as to how the semantic web would help, I think your suggestion gives the evidence. Google is automatically extracting info based on its semantics, ie if it looks like an address, attach that semantic info to it, and utilise it.
This also goes a long way to back up your point that the semantic web will probably start if we can do it automatically. Then I could see manual initiatives starting along the lines of wikipedia: someone has a cool way of adding semantic information via the crowdsource. That is when we will have a breakthrough in the SW
And as to how the semantic web would help, I think your suggestion gives the evidence. Google is automatically extracting info based on its semantics, ie if it looks like an address, attach that semantic info to it, and utilise it.
This also goes a long way to back up your point that the semantic web will probably start if we can do it automatically. Then I could see manual initiatives starting along the lines of wikipedia: someone has a cool way of adding semantic information via the crowdsource. That is when we will have a breakthrough in the SW
Maybe it's time people in your country start looking at other search engines :-)
Why is Google only number 3? And what does it matter if another search engine is more popular? Just use the best one. Of course maybe others are better optimized for your country, but somehow I doubt it. Also, how many searches are country specific anyway?
Why is Google only number 3? And what does it matter if another search engine is more popular? Just use the best one. Of course maybe others are better optimized for your country, but somehow I doubt it. Also, how many searches are country specific anyway?
Google readily admits that they can't do local well and would love to do it well. They also can't do real-time. Those are both specialties of the SW.
The more you argue, the more you show your lack of understanding both with the state of the art and what's possible.
The more you argue, the more you show your lack of understanding both with the state of the art and what's possible.
But the semantics of searching on a map is not the "semantic web" that people like Tim Berners-Lee are talking about.
It people out there can't be bothered to search google maps, is there some particular reason they'd instantiate a reasoning engine on a semantic graph to find a local babysitter?
It people out there can't be bothered to search google maps, is there some particular reason they'd instantiate a reasoning engine on a semantic graph to find a local babysitter?
Oh yes it is. In his first TED speech his exact example is naming a point on a map.
Demonstrate a working, genuinely useful semantic system and the world will believe you. Until then, the SW is a monumental waste of everybody's time.
this is the semantic web http://en.wikipedia.org/wiki/Semantic_Web
Searching google maps is not the semantic web no matter how you try to twist the term. You can't both claim the SW already exists and doesn't exist at the same time. The fact that I can already do something that the SW claims it will be able to do, in 15 or 20 years, already demonstrates that the principle isn't necessary and is already obsolete.
this is the semantic web http://en.wikipedia.org/wiki/Semantic_Web
Searching google maps is not the semantic web no matter how you try to twist the term. You can't both claim the SW already exists and doesn't exist at the same time. The fact that I can already do something that the SW claims it will be able to do, in 15 or 20 years, already demonstrates that the principle isn't necessary and is already obsolete.
Maybe it's time people in your country start looking at other search engines :-)
Why is Google only number 3? And what does it matter if another search engine is more popular? Just use the best one. Of course maybe others are better optimized for your country, but somehow I doubt it.
Why is Google only number 3? And what does it matter if another search engine is more popular? Just use the best one. Of course maybe others are better optimized for your country, but somehow I doubt it.
> Why didn't she use the web? Because she can't.
Please describe why she can't?
I shall now demonstrate why she can. Here's an example in the states.
http://www.sitters.com/default.htm
That took me approximately 1 second to find (give or take some insignificant fraction of a second).
Or this.
http://www.expats.cz/prague/t-284477.html
Wow, incredible! You want another couple hundred sources? If I could speak Czech I could probably find the same.
There is absolutely nothing proposed with the SW that would have improved this. And in fact, an unranked, unattributed list of babysitters returned from some semantic datastore no matter how large (without attributable trust signals) will change that.
Your example, would have required her to enter her babysitting availability in the SW instead of on some web site. Doesn't really make a difference, if it's crawlable, it's already retrievable.
> Until someone makes a Czech babysitting database, she can't. Because no one would ever find it.
Here, http://www.bibo-hlidani-deti.cz/ wow, impossible.
Total retrieval time, in two languages, one of which I don't speak, 30 seconds. I can only imagine what I could find if I spoke Czech. Does Google not operate in your country?
Is there some reason that building, maintaining and populating a "Czech babysitting database" is fundamentally different from populating the SW?
>That's data. So why don't they put it on the web as text? Because no one would ever find it. Text doesn't work that way.
The SW isn't some magical self populating entity made out of wishes. It would take an effort worth the cost of buying France to build a semantic web of the size and information density of the current Web.
Furthermore, nobody has ever demonstrated why it would be better. Statement like "Whoever fixes this, changes it all. They take out Google. They take out Monster.com. They take out Match.com. They take out Facebook. Whoever can do this will change everything." are not only built on a completely incorrect premise -- searching large volumes of indexed text and providing a ranked relevant result set is impossible -- they aren't even built on the ability to point to a working, large scale semantic graph with the appropriate reasoning engines (which in an of themselves are barely in their most basic infancy)...not one presently exists, and the smaller scale examples haven't yielded much of anything valuable.
And in many obvious ways it would be worse. People can put effort into a web page to send signals about quality, other sites and sources can rank a page as trustworthy or not. A list of "babysitters in the Czech Republic" offers none of that. And reasoning out that information across an infinitely large graph could take hours.
I've worked with large semantic systems, with billions of entities and tens of billions of relationships, and almost universally they provide almost no additional value over traditional search, yet perform much more poorly.
Please describe why she can't?
I shall now demonstrate why she can. Here's an example in the states.
http://www.sitters.com/default.htm
That took me approximately 1 second to find (give or take some insignificant fraction of a second).
Or this.
http://www.expats.cz/prague/t-284477.html
Wow, incredible! You want another couple hundred sources? If I could speak Czech I could probably find the same.
There is absolutely nothing proposed with the SW that would have improved this. And in fact, an unranked, unattributed list of babysitters returned from some semantic datastore no matter how large (without attributable trust signals) will change that.
Your example, would have required her to enter her babysitting availability in the SW instead of on some web site. Doesn't really make a difference, if it's crawlable, it's already retrievable.
> Until someone makes a Czech babysitting database, she can't. Because no one would ever find it.
Here, http://www.bibo-hlidani-deti.cz/ wow, impossible.
Total retrieval time, in two languages, one of which I don't speak, 30 seconds. I can only imagine what I could find if I spoke Czech. Does Google not operate in your country?
Is there some reason that building, maintaining and populating a "Czech babysitting database" is fundamentally different from populating the SW?
>That's data. So why don't they put it on the web as text? Because no one would ever find it. Text doesn't work that way.
The SW isn't some magical self populating entity made out of wishes. It would take an effort worth the cost of buying France to build a semantic web of the size and information density of the current Web.
Furthermore, nobody has ever demonstrated why it would be better. Statement like "Whoever fixes this, changes it all. They take out Google. They take out Monster.com. They take out Match.com. They take out Facebook. Whoever can do this will change everything." are not only built on a completely incorrect premise -- searching large volumes of indexed text and providing a ranked relevant result set is impossible -- they aren't even built on the ability to point to a working, large scale semantic graph with the appropriate reasoning engines (which in an of themselves are barely in their most basic infancy)...not one presently exists, and the smaller scale examples haven't yielded much of anything valuable.
And in many obvious ways it would be worse. People can put effort into a web page to send signals about quality, other sites and sources can rank a page as trustworthy or not. A list of "babysitters in the Czech Republic" offers none of that. And reasoning out that information across an infinitely large graph could take hours.
I've worked with large semantic systems, with billions of entities and tens of billions of relationships, and almost universally they provide almost no additional value over traditional search, yet perform much more poorly.
What is the incentive for people to do it? It is work, so they should get something in return.
Maybe SEO for Google, ie if Google said "we'll improve your rank dramatically if you included semantic data". But even then - it would be an enormous amount of work to update all (or the majority of web sites).
Maybe SEO for Google, ie if Google said "we'll improve your rank dramatically if you included semantic data". But even then - it would be an enormous amount of work to update all (or the majority of web sites).
Why would they do it? To be found.
There are over a trillion web sites. Why did people put in so much work? Because they had something that they wanted to be found/seen/known by someone else.
The semantic web, in the form that it will come in, won't be nearly that much work. Instead of each business creating a web site, they'll just fill out the data of what makes them what they are and what makes them unique.
There are a trillion textual web sites, nearly all created in the last 15 years, most of which don't need to exist. My prediction: Very soon, they won't.
Edit: I'm not talking about current initiatives like LinkedData. If the SW would only come like that, I would agree with you.
There are over a trillion web sites. Why did people put in so much work? Because they had something that they wanted to be found/seen/known by someone else.
The semantic web, in the form that it will come in, won't be nearly that much work. Instead of each business creating a web site, they'll just fill out the data of what makes them what they are and what makes them unique.
There are a trillion textual web sites, nearly all created in the last 15 years, most of which don't need to exist. My prediction: Very soon, they won't.
Edit: I'm not talking about current initiatives like LinkedData. If the SW would only come like that, I would agree with you.
I think your reasoning about why people make web pages is error-prone.
Just another node on a semantic graph is not what anybody wants to be, otherwise the web would be populated with HTML 1.0 pages. It's not like making one of those is much harder than filling out a semantic form.
> most of which don't need to exist.
Who says that? I think my page on mauve socks should certainly exist. There's an awful lot of stuff like that on the web. And the SW is no more useful for finding that than search is.
Just another node on a semantic graph is not what anybody wants to be, otherwise the web would be populated with HTML 1.0 pages. It's not like making one of those is much harder than filling out a semantic form.
> most of which don't need to exist.
Who says that? I think my page on mauve socks should certainly exist. There's an awful lot of stuff like that on the web. And the SW is no more useful for finding that than search is.
Google actually has rolled out a small step towards this. They recognize microformats and other bits of structured data, and use them to do things like supply the number of stars for a review or to populate friend lists for social networking sites.
Their official jargon is "Rich Snippets", and they have a couple posts on it at http://googlewebmastercentral.blogspot.com/2009/05/introduci... and http://googlewebmastercentral.blogspot.com/2009/10/help-us-m... .
Their official jargon is "Rich Snippets", and they have a couple posts on it at http://googlewebmastercentral.blogspot.com/2009/05/introduci... and http://googlewebmastercentral.blogspot.com/2009/10/help-us-m... .
Maybe SEO for Google
And, of course, the "semantic web" will be able to disregard bad data (published out of malice, incompetence or bitrot) because it ... hmm ... because ... hmm, er ... ah, well. They must have it covered.
(I'm somewhat reminded of this: http://en.wikipedia.org/wiki/Expert_system , only with the maintenance problems squared or cubed.)
P.S. The RethinkDB guys (Slava, etc.) are chasing an interesting problem (DBs on SSD). I submitted a link to their blog: http://news.ycombinator.com/item?id=1337391
And, of course, the "semantic web" will be able to disregard bad data (published out of malice, incompetence or bitrot) because it ... hmm ... because ... hmm, er ... ah, well. They must have it covered.
(I'm somewhat reminded of this: http://en.wikipedia.org/wiki/Expert_system , only with the maintenance problems squared or cubed.)
P.S. The RethinkDB guys (Slava, etc.) are chasing an interesting problem (DBs on SSD). I submitted a link to their blog: http://news.ycombinator.com/item?id=1337391
I don't understand the criticism that people have to markup all their data. As I understand the semantic web, only programmers and certain other skilled users were expected to knowingly do that. (Because the semantic web is more about providing interop between programs.)
So, when I type a date into a box, that date can later be represented as a departure date. Some programmer already constrained and validated my input, and stored it in a DB; I don't have to do anything unusual. Later, semantic web clients will request my departure date (among other things); then it's pulled out of the DB or cache, and formatted in the semantic web format which they expect.
Do I misunderstand?
So, when I type a date into a box, that date can later be represented as a departure date. Some programmer already constrained and validated my input, and stored it in a DB; I don't have to do anything unusual. Later, semantic web clients will request my departure date (among other things); then it's pulled out of the DB or cache, and formatted in the semantic web format which they expect.
Do I misunderstand?
Not at all; you got it.
Others here, though, seem to have either no idea, or worse, a deeply misguided idea of what the Semantic Web is. I find I'm not arguing for the Semantic Web but against a lot of deeply ingrained misconceptions.
Others here, though, seem to have either no idea, or worse, a deeply misguided idea of what the Semantic Web is. I find I'm not arguing for the Semantic Web but against a lot of deeply ingrained misconceptions.
I don't think anybody is arguing against semantic information. By "semantic web" I primarily think about the effort of creating new markup languages for the web to enable people to mark their data with semantic information that is machine readable. The problem then is, how do you get people to add that extra information.
If you can extract information automatically, then you don't need to explicitly mark the data to begin with, so in my opinion it is not the "semantic web" people have been talking about.
Of course there was meta-information from day one, that is, web forms had fields for dates and names and in the database there were columns called date and name. That is Web 1.0, not semantic web.
If you can extract information automatically, then you don't need to explicitly mark the data to begin with, so in my opinion it is not the "semantic web" people have been talking about.
Of course there was meta-information from day one, that is, web forms had fields for dates and names and in the database there were columns called date and name. That is Web 1.0, not semantic web.
That would only be a partial solution, I think.
Have you seen this classic piece from 2001?
http://www.well.com/~doctorow/metacrap.htm
It argues quite convincingly why semantic web is useless.
It argues quite convincingly why semantic web is useless.
The author's arguments are: people are lazy, stupid and liars.
Dear author, fuck you. We, the "common man" have been denigrated and belittled since our existence began. Yet we made the web. We made Wikipedia. And we'll do it again.
Dear author, fuck you. We, the "common man" have been denigrated and belittled since our existence began. Yet we made the web. We made Wikipedia. And we'll do it again.
I have to point out that the author's essay was from 2001. (About 7 months after wikipedia registered its domain name) So on the one hand, the author's insights are a little outdated. On the other hand, sad to say, most of what he points out is almost timeless. People are lazy and liars. Some people (spammers) are industrious and liars.
I think his point about 'There's more than one way to describe something' is the real crux of the issue. I call it art, you call it crap, where does it fit into the SW?
My opinion is that we are at the beginnings of even trying to figure out what it is we want. Semantic web seems to shine a light on a possible better way, but I think we are just starting out. I think SW will be one part of the internet just like the web is just one part.
I think his point about 'There's more than one way to describe something' is the real crux of the issue. I call it art, you call it crap, where does it fit into the SW?
My opinion is that we are at the beginnings of even trying to figure out what it is we want. Semantic web seems to shine a light on a possible better way, but I think we are just starting out. I think SW will be one part of the internet just like the web is just one part.
Sure, you're right. People will always spam and lie, etc, but I don't see that as a blocker for the development of a SW as much as it was a blocker for the web itself.
I think there's already so much in terms of the data that's in social sites, dating sites, real estate sites, etc. And it's all fine. There's also a ton in RDF triples already in LinkedData, etc. And that stuff is fine too. I really think it's not as bad as our fears.
But I'm noticing I'm not getting much love on this thread. Very interesting. Chris Dixon states in the (OP) video that the Semantic Web has turned into a dirty word. I can see that's true now. Fascinating.
I think there's already so much in terms of the data that's in social sites, dating sites, real estate sites, etc. And it's all fine. There's also a ton in RDF triples already in LinkedData, etc. And that stuff is fine too. I really think it's not as bad as our fears.
But I'm noticing I'm not getting much love on this thread. Very interesting. Chris Dixon states in the (OP) video that the Semantic Web has turned into a dirty word. I can see that's true now. Fascinating.
I just shared the video on Twitter and LinkedIn. Great, informative video, if you ask me. It depicts both points of view.
However, I do agree that the word "semantic" has become too much of a buzzword and most people aren't really aware of it's true meaning.
However, I do agree that the word "semantic" has become too much of a buzzword and most people aren't really aware of it's true meaning.
Well done short documentary on the promise and problems of the semantic web. A lot of familiar names interviewed on the topic. Berners-Lee, Shirky, etc.
I was shocked; I think it's everything.
The Semantic Web is the big nut. Just because we don't have a way to crack it yet, doesn't diminish it in terms of its importance. This video helps to hint at why that is and what's at stake.