Rel=next and rel=prev - how Google might be killing a spec(onderhond.com)
onderhond.com
Rel=next and rel=prev - how Google might be killing a spec
http://www.onderhond.com/blog/work/rel-next-prev-google
5 comments
> I think Google is right here. If you mark something as a sequence, it is reasonable to presume that the start of the sequence is the normal entry point, in this case the front page of your blog.
I disagree. Just because the rel="prev"/rel="next" commands provide the ability to browse forward and backward through a series of documents doesn't mean that the only possible entry point is the beginning of the series. It doesn't mean that there even is a single well-defined first entry in the series. There's nothing to stop a series of links connecting into a closed loop, or to stop a single page from having multiple rel="prev" links because it appears in multiple series, or multiple rel="next" links because a previously linear series bifurcates and the user has multiple choices.
I disagree. Just because the rel="prev"/rel="next" commands provide the ability to browse forward and backward through a series of documents doesn't mean that the only possible entry point is the beginning of the series. It doesn't mean that there even is a single well-defined first entry in the series. There's nothing to stop a series of links connecting into a closed loop, or to stop a single page from having multiple rel="prev" links because it appears in multiple series, or multiple rel="next" links because a previously linear series bifurcates and the user has multiple choices.
>> Generally ordered sequences are designed to be read from the start
The spec makes no such claims, it just talks about sequences of documents in a very generic way. If the spec is not clear on this, it should be changed first.
And in a search context, I definitely don't agree with your interpretation of "sequence". The fact that my hit is part of a sequence is secondary to the fact that I just want to see a hit on my search query.
Mind you that I don't necessarily disagree with your nuances on the various use cases, but the current spec is clearly written to be very generic and should be interpreted as such.
The spec makes no such claims, it just talks about sequences of documents in a very generic way. If the spec is not clear on this, it should be changed first.
And in a search context, I definitely don't agree with your interpretation of "sequence". The fact that my hit is part of a sequence is secondary to the fact that I just want to see a hit on my search query.
Mind you that I don't necessarily disagree with your nuances on the various use cases, but the current spec is clearly written to be very generic and should be interpreted as such.
If I'm looking for one paragraph in the middle of a long sequence of pages I want to be taken to the page that paragraph is on, not the first page.
A web-log is, for a lot of people, a historical (chronological) record. For these cases rel=next/prev makes sense to semantically create a relationship between individual post.
A web-log is, for a lot of people, a historical (chronological) record. For these cases rel=next/prev makes sense to semantically create a relationship between individual post.
> I think Google is right here. If you mark something as a sequence, it is reasonable to presume that the start of the sequence is the normal entry point, in this case the front page of your blog.
Imagine applying your rule to Google Book Search. Google found the word "fizbit" on page 2031 of a book. I click on the link and Google takes me to the FIRST PAGE of the book. So I get to hunt through the book page by page until I find that word.
The same thing applies to web pages of all kinds. Here's my maxim: if Google has a result for your query to some term, then clicking on that result better take you to a page with the term on it. Otherwise Google's result was a waste of your time.
Imagine applying your rule to Google Book Search. Google found the word "fizbit" on page 2031 of a book. I click on the link and Google takes me to the FIRST PAGE of the book. So I get to hunt through the book page by page until I find that word.
The same thing applies to web pages of all kinds. Here's my maxim: if Google has a result for your query to some term, then clicking on that result better take you to a page with the term on it. Otherwise Google's result was a waste of your time.
> I think Google is right here. If you mark something as a sequence, it is reasonable to presume that the start of the sequence is the normal entry point, in this case the front page of your blog.
For me, if I click on a URL I want to go to that URL. I don't want the search engine or the browser deciding that I'm better off going to some other URL. Google should provide the information, let me be the driver. If this gets too annoying I'll just switch to ddg.gg.
I agree that the author is misusing prev/next if he plans to put them on ever article in his blog. Using prefetch for similar articles makes sense though.
For me, if I click on a URL I want to go to that URL. I don't want the search engine or the browser deciding that I'm better off going to some other URL. Google should provide the information, let me be the driver. If this gets too annoying I'll just switch to ddg.gg.
I agree that the author is misusing prev/next if he plans to put them on ever article in his blog. Using prefetch for similar articles makes sense though.
> For me, if I click on a URL I want to go to that URL.
I don't think that was ever up for debate. This is about which pages appear on search results, so you might see the url of the first page of an article instead of the last page.
This all seems overblown. If you add metadata, expect it to be used!
I don't think that was ever up for debate. This is about which pages appear on search results, so you might see the url of the first page of an article instead of the last page.
This all seems overblown. If you add metadata, expect it to be used!
This is not my interpretation of the Google blog post[1]. If your search result is on page 3 they'll send you to page 1. This is undesirable to me. Please call me out if this is an incorrect interpretation.
[1]http://googlewebmastercentral.blogspot.com/2011/09/paginatio...
[1]http://googlewebmastercentral.blogspot.com/2011/09/paginatio...
They say they may. It may also depend how similar the content on the pages is, and how specific your search term is. Eg if there are 10 pages on a disease, one is about symptoms, and your search term is about symptoms you probably want that page, if it is a general search you probably want the start.
Thats my interpretation of them, how it works in actual use we shall see.
Thats my interpretation of them, how it works in actual use we shall see.
Sorry, the difficulty seems to be what "search result" means. If you search for "foo", and that term happens to be on page 3, then Google may produce page 1 as the search result. Just like at the moment, sometimes the term you search for only appears on the pages which link to the search result, not on the search result itself.
Sometimes the term does not appear at all, for example if I search for "batman begins sequel" the first result is: http://www.imdb.com/title/tt0468569/ That page does not have the term "sequel" on it, but it is still the first search result. And if you view the cached version, the only word highlighted is "batman".
Very specifically, my post above addresses the comment that Google might display one url, then take you to another. That seems unlikely. Instead, you will only see the url of the page that is the search result, which may be page 1.
Sometimes the term does not appear at all, for example if I search for "batman begins sequel" the first result is: http://www.imdb.com/title/tt0468569/ That page does not have the term "sequel" on it, but it is still the first search result. And if you view the cached version, the only word highlighted is "batman".
Very specifically, my post above addresses the comment that Google might display one url, then take you to another. That seems unlikely. Instead, you will only see the url of the page that is the search result, which may be page 1.
I agree with your reasoning completely, but I also enjoy going from one distinct article to the next with a fast-forward command or gesture. Is there a way to enable that style of browsing using something other than rel-next?
Opera has been guessing which link to use as "next" for a long time, the feature is called "fast forward".
...and I use rel-next to enable fast forward. I could be ignorant of other methods though.
If rel=next is missing, Opera guesses based on link label. There used to be a fast_forward.ini file with list of all recognized names ("Next page", "Proceed", "Forward", etc.).
Fast forward works in google search results, which doesn't use rel="next". I suspect that Opera does it heuristically, so if it sees a link with the text "next" it will use that.
[deleted]
If Google wants to find first document in series, they should have RTFM. There's relationship that specifies exactly this — no need to twist meaning of prev/next:
> Start: Refers to the first document in a collection of documents. This link type tells search engines which document is considered by the author to be the starting point of the collection.
http://www.w3.org/TR/html4/types.html#type-links
> Start: Refers to the first document in a collection of documents. This link type tells search engines which document is considered by the author to be the starting point of the collection.
http://www.w3.org/TR/html4/types.html#type-links
The author is selling google way short here: "It's a little scary to think that one company (~85% of the world population searches the web with Google) can make such a trifle assumption and make a simple, clear cut spec like this virtually unusable." In short, I believe the author is making an unfounded and probably wrong assumption when he says that Google is making an assumption.
It seems to me from reading Google's blog post on the subject (http://googlewebmastercentral.blogspot.com/2011/09/paginatio...) that they're basing their behavior on what they believe to be preferential to their users based on the data they've collected. Here's an example: "Because view-all pages are most commonly preferred by searchers, we do our best to surface this version..." They don't say "we are assuming view-all pages are preferred," they actually know that they are.
It's true that Google's post doesn't say something similar about why they choose to surface the first page in a sequence, but I think their reputation suggests that they have numbers to support the idea. In any case, I wouldn't assume that their behavior is based on an assumption.
It seems to me from reading Google's blog post on the subject (http://googlewebmastercentral.blogspot.com/2011/09/paginatio...) that they're basing their behavior on what they believe to be preferential to their users based on the data they've collected. Here's an example: "Because view-all pages are most commonly preferred by searchers, we do our best to surface this version..." They don't say "we are assuming view-all pages are preferred," they actually know that they are.
It's true that Google's post doesn't say something similar about why they choose to surface the first page in a sequence, but I think their reputation suggests that they have numbers to support the idea. In any case, I wouldn't assume that their behavior is based on an assumption.
Google bases their actions on dry information harvested from people. Human science does not exist without statistics, and statistics requires you to make assumptions. Google can analyze me till kingdom come, it cannot correctly predict what I prefer at a certain moment in time. And as for big companies making documented assumptions, check what Facebook rolled out today in their news feed.
In situations where user preference is key, the only option is to leave the decision to the individual user. I don't mind if Google implements this and offers it as a service to their users, but deciding based on dry statistics is never good, it leaves people frustrated and makes it impossible for authors to properly implement a spec like this.
In situations where user preference is key, the only option is to leave the decision to the individual user. I don't mind if Google implements this and offers it as a service to their users, but deciding based on dry statistics is never good, it leaves people frustrated and makes it impossible for authors to properly implement a spec like this.
I work on localsearch ranking at Google, but this I don't know anything about Google's use of this information in web search, nor am I speaking for Google.
This is about the philosophy of how we approach problems in ranking.
The appropriate question is not, "is this going to work everytime?". If that was the question, basically no one would make progress in web search ever. The web is littered with randomness and corner cases. The question is, "is this a a generally good heuristic, will it help more than it hurts, and can we detect when the heuristic fails badly?" I am guessing that if this prev/next information is used, it will be to boost the front page of an article that has SOME support for the query over the second or third or n'th page of an article that has more, but not exceptionally more support for the query.
Don't think of this as a hard if, but more like an additional weight in a summed comparison that biases towards the first page over later pages in a sequence.
This is about the philosophy of how we approach problems in ranking.
The appropriate question is not, "is this going to work everytime?". If that was the question, basically no one would make progress in web search ever. The web is littered with randomness and corner cases. The question is, "is this a a generally good heuristic, will it help more than it hurts, and can we detect when the heuristic fails badly?" I am guessing that if this prev/next information is used, it will be to boost the front page of an article that has SOME support for the query over the second or third or n'th page of an article that has more, but not exceptionally more support for the query.
Don't think of this as a hard if, but more like an additional weight in a summed comparison that biases towards the first page over later pages in a sequence.
Yes, my thoughts too. The author pulled this quote from a Google document but is there any evidence The Google Search Engine is even implementing this behavior in a negative way?
I doubt there’s a company in the world that doesn’t say it does X in some document but do Y in reality. That’s the problem with documentation: it rots.
I doubt there’s a company in the world that doesn’t say it does X in some document but do Y in reality. That’s the problem with documentation: it rots.
Google seems to be interpreting prev/next wrongly. Suppose you are searching for the 2nd episode of your preferred TV series. Episodes are sibling resources, and this concept might be represented with prev/next. However you would expect to get to the right page, not the 1st episode, when you click on the link.
What Google could do, is showing a separate link to the first item in the series, and let the user to decide. Later monitoring the clicks it could provide a customised experience per user.
Isn't this more of a problem with rel="start" than rel="next" and rel="prev"?
I think Google is right here. If you mark something as a sequence, it is reasonable to presume that the start of the sequence is the normal entry point, in this case the front page of your blog.
Personally, the obsession about date order in blogs is overdone. It is not a book, there are no real dependencies between chapters. My blog is not a sequence (unless I happen to say write a three part article on some subject, where there would be a short sequence here). Sure it has a date order, but semantically thats not very important. I am not self important enough to believe anyone reads it in order. If people want to read by date, you can provide a date index or calendar view, but that does not have to imply a rel=next/prev sequence on that view.