Hashify: what becomes possible when one is able to store documents in URLs?(hashify.me)
http://hashify.me/IyBIYXNoaWZ5CgpIYXNoaWZ5IGRvZXMgbm90IHNvbHZlIGEgcHJvYmxlbSwgaXQgcG9zZXMgYSBxdWVzdGlvbjogX3doYXQgYmVjb21lcyBwb3NzaWJsZSB3aGVuIG9uZSBpcyBhYmxlIHRvIHN0b3JlICoqZW50aXJlIGRvY3VtZW50cyoqIGluIFVSTHM/XwoKIyMgRG9jdW1lbnQg4oaUIFVSTAoKSGFzaGlmeSBpcyBkaWZmZXJlbnQgZnJvbSB2aXJ0dWFsbHkgZXZlcnkgb3RoZXIgc2l0ZSBvbiB0aGUgV2ViIGluIHRoYXQgKipldmVyeSBVUkwgY29udGFpbnMgdGhlIGNvbXBsZXRlIGNvbnRlbnRzIG9mIHRoZSBwYWdlKiouCgpUaGUgYWRkcmVzcyBiYXIgdXBkYXRlcyB3aXRoIGVhY2gga2V5c3Ryb2tlIGFzIG9uZSB0eXBlcyBpbnRvIHRoZSBlZGl0b3IuCgojIyMgQmFzZTY0IGVuY29kaW5nCgpPbmx5IGEgdGlueSBmcmFjdGlvbiBvZiBhbGwgVW5pY29kZSBjaGFyYWN0ZXJzIGFyZSBhbGxvd2VkIHVuZXNjYXBlZCBpbiBhIFVSTC4gSGFzaGlmeSB1c2VzIFtCYXNlNjRdWzFdIGVuY29kaW5nIHRvIGNvbnZlcnQgVW5pY29kZSBpbnB1dCB0byBBU0NJSSBvdXRwdXQgc2FmZSBmb3IgaW5jbHVzaW9uIGluIFVSTHMuCgpUaGlzIHRyYW5zbGF0aW9uIGlzIGEgdHdvLXN0ZXAgcHJvY2VzczogW1VuaWNvZGUgdG8gVVRGLTggY29udmVyc2lvbl1bMl0gYXMgb3V0bGluZWQgYnkgSm9oYW4gU3VuZHN0csO2bSwgZm9sbG93ZWQgYnkgYmluYXJ5IHRvIEFTQ0lJIGNvbnZlcnNpb24gdmlhIFtgd2luZG93LmJ0b2FgXVszXS4KCiMjIyMgRW5jb2RpbmcKCiAgICA+IHVuZXNjYXBlKGVuY29kZVVSSUNvbXBvbmVudCgnw6dhIHZhPycpKQogICAgIsODwqdhIHZhPyIKICAgID4gYnRvYSh1bmVzY2FwZShlbmNvZGVVUklDb21wb25lbnQoJ8OnYSB2YT8nKSkpCiAgICAidzZkaElIWmhQdz09IgoKIyMjIyBEZWNvZGluZwoKICAgID4gYXRvYigndzZkaElIWmhQdz09JykKICAgICLDg8KnYSB2YT8iCiAgICA+IGRlY29kZVVSSUNvbXBvbmVudChlc2NhcGUoYXRvYigndzZkaElIWmhQdz09JykpKQogICAgIsOnYSB2YT8iCgojIyBVUkwgc2hvcnRlbmluZwoKU3RvcmluZyBhIGRvY3VtZW50IGluIGEgVVJMIGlzIG5pZnR5LCBidXQgbm90IHRlcnJpYmx5IHByYWN0aWNhbC4gSGFzaGlmeSB1c2VzIHRoZSBbYml0Lmx5IEFQSV1bNF0gdG8gc2hvcnRlbiBVUkxzIGZyb20gYXMgbWFueSBhcyAzMCwwMDAgY2hhcmFjdGVycyB0byBqdXN0IDIwIG9yIHNvLiBJbiBlc3NlbmNlLCBiaXQubHkgYWN0cyBhcyBhIGRvY3VtZW50IHN0b3JlIQoKIyMjIFVSTCBsZW5ndGggbGltaXQKCldoaWxlIHRoZSBIVFRQIHNwZWNpZmljYXRpb24gZG9lcyBub3QgZGVmaW5lIGFuIHVwcGVyIGxpbWl0IG9uIHRoZSBsZW5ndGggb2YgYSBVUkwgdGhhdCBhIHVzZXIgYWdlbnQgc2hvdWxkIGFjY2VwdCwgYml0Lmx5IGltcG9zZXMgYSAyMDQ4LWNoYXJhY3RlciBsaW1pdC4gVGhpcyBpcyBzdWZmaWNpZW50IGluIHRoZSBtYWpvcml0eSBvZiBjYXNlcy4KCkZvciBsb2
31 comments
Git.
-
It may be of interest to view this duality as an analog to the duality of location addressing (iterative) vs value addressing (functional) in context of memory mangers. The general (hand wavy as of now) idea is a distributed memory system with a functional front-end (e.g. Scala/Haskell).
"A New Way to look at Networking" http://www.youtube.com/watch?v=8Z685OF-PS8
I'm not quite sure if URN is exactly right for the hash thing either, given that it both fails to unify things which humans would probably assign the same URN to, such as two image files of the same picture using different encodings, and it has the theoretical chance of assigning the same hash to two entirely different things.
Names: universally unique and fully scoping the life-cycle of the (logical) object. 1:1.
Identifiers: unique in context of an authority with a life-cycle that is maximally (but not necessarily) bounded by the life-cycle of the named entity (and of course, the authority that assigns it). e.g. http://www.ssa.gov/history/ssn/geocard.html is the authority that issues SSN identifiers. An entity can potentially have multiple such identifiers. 1:N
Locations: The location of an image or representation of the entity. 1:N (e.g. CDNs)
* Global uniqueness: The same URN will never be assigned to two different resources. ((the encoding would be part of the URN))
* Independence: It is solely the responsibility of a name issuing authority to determine the conditions under which it will issue a name. ((a URN wouldn't necessarily be a hash of the resource in question))
The second point makes it pretty clear that the assignment of URN's would be done by some authoritative parties, which makes sense if you think that in their initial view URN's would have been useful in linking citations, references; for research papers. Just that the Internet has far time ago branched from that scope.
http://code.google.com/p/dijjer/
In a single link for a file, it can contain multiple hashes for multiple means of retrieval.
Some years ago a friend and I wrote http://notamap.com, a very similar idea for sharing/storing/embed geotagged notes fully encoded on a URL, without having to rely on a server. Looking at it now I wish we had not put all the crazy animations. Maybe I should recover it and simplify the UI.
I just can't see the gain here. You need a server to distribute the URLs in any case. You are just moving the data from the server that served the document to the server that servers the URLs. It is still the same data, just in different form.
For this to work you need to be able to recover the document from the URL locally.
How about saving the document?
One step further in this direction would just mean "including all the data for one or more documents on one page". You've just invented...a document, possibly with more than one media type.
Think about it this way: You have a 10K document which contains 200 bytes of link data; a URI to another hypertext document that is 5K in size...vs...you have a 15K document which contains 5K of "link data"; the other document.
With those 2 things, everything should be covered.
* TinyURL 65,536 characters and probably more, but requests timed out; there isn’t an explicit limit apparently
* Bit.ly 2000 characters.
* Is.Gd 2000 characters.
* Twurl.nl 255 characters.
This was 2.5 years ago, not sure how much of these have changed (other than bit.ly, which the linked article confirms is 2048, probably the same as when I tested it).
http://softwareas.com/the-url-shortener-as-a-cloud-database
Interesting, but I can't think of any practical application, apart from the service provider not having to worry about storage (maybe that's key ... more thinking needed).
This way the whole tool can be 100% client-side javascript, without a need for any back-end.
The project later moved to http://glsl.heroku.com/ with an app-driven gallery, and that particular feature went away. I think that is a pretty natural evolution of any such idea, so I'm not convinced of hashify's logevity, but hey, simple sometimes is really enough.
Trouble is, there might be a DNS-like system needed to match hashify URLs to more human-readable strings (or a way for existing DNS to resolve to hashify style URLs).
Neat idea.
"Storing a document in a URL is nifty, but not terribly
practical. Hashify uses the [bit.ly API][4] to shorten
URLs from as many as 30,000 characters to just 20 or so.
In essence, bit.ly acts as a document store!"
http://tinyurl.com/3n6h8pxmaybe switch to goo.gl?
But what happens if bit.ly just say "Sorry incoming URLS (long urls) can only be a maximum of 1500 chars"?
While the HTTP specification does not define an upper limit
on the length of a URL that a user agent should accept,
bit.ly imposes a 2048-character limit.Think about this like a PDF where stuff is embedded instead of in separate files.
But URL shortening services are a public good, and hacking one to be your personal cloud storage platform is kind of a dick move.
What has been will be again,
what has been done will be done again;
there is nothing new under the sun.
- Ecclesiates 1:9http://www.semicomplete.com/projects/keynav/
There's also snappy http://code.google.com/p/snappy/
Plaintext, the past, present and future.
There's also several ways to obscure the impact of SOPA on the URL shortening anyway. For instance, if several services use the same hash algorithm for representing URLs, they can be used interchangeably (if you post the URL to all of them). Further, you can always set up your own temporary shortening service as well.
Altered to convey another point. Naturally, it would be quite difficult to "embed" a feature length movie into a single url, but if one was to split the file into chunks like torrent transferring does, or simply a multi-part rar like newsgroups still do, it enables each chunk to be more manageable.
I do agree with you though, but I think the reason that a service like this if changed in such a way to be user-friendly for file sharing, not just document sharing, would be able to get around a lot of the pitfalls a torrent tracker (for example) would have if it's DNS lookup was blocked (which aren't many) is due to the simple fact that SOPA is written in a way that assumes all IP addresses and DNS names are statically tied together and slow to alter, not that I can have a new domain name in a matter of minutes that resolves to my existing server. Even more so if the final URL hash was nothing more than a common and known algorithm, like base64, that one could easily plug into a basic desktop app and get the same result.
Is it just me or does anyone else also just back away whenever there is a project which turns nouns into verbs with ify? Spotify is a sockpuppet of
The tricky part with that system would be that you'd also need some new mechanism to retrieve the files. Instead of the regular WWW stack, you'd need something like a massive distributed hash table that could handle massive distributed querying and transferring the hashed files. Many P2P file sharing systems are already doing this, but a sparse collection of end-user machines containing a few hashed files each isn't a very efficient service cloud. If every ISP had this sort of thing in their service stack or if Amazon and Google decided to run the service, all of them dynamically caching documents in greater demand in more nodes, things might look very different.
This would mean that very old hypertext documents would still be trivially readable with working links, as long as a few copies of the page documents were still hashed somewhere, even if the original hosting servers were long gone. It would also make it easy to do distributed page caching, so that pages that get a sudden large influx of traffic wouldn't create massive load on a single server.
On the other hand, any sort of news sites where the contents of the URL are expected to change wouldn't work, nor would URLs expected to point to a latest version of a document instead of the one at the time of linking. Once the hash URL was out, no revision to the hashed document visible from following the URL would be possible without some additional protocol layer. The URL strings would also be opaque to humans and too long and random to be committed to memory or typed by hand. The web would probably need to be somehow split into human-readable URLs for dynamic pages and hash URLs for the static pieces of content served by those pages.
I'm probably reinventing the wheel here, and someone's already worked out a more thought out version of this idea.