Any specific sites? Happy to spin one of these up for you focused on Web3. These work the best when the search engine creator has domain expertise to showcase the best sources.
I could try to find some sources, but my guess at the best Web3 sites would probably miss the mark.
That's great! And I think the results even improve a little more if you just search 'distributed model parallel' without 'pytorch' [0]
So still some work to do, but I've personally found this type of search engine really valuable for some of my interests where I am able to curate the sites that might not normally rank, but I know are higher quality.
Yeah, that's a good point. As new ones get built, we will need a way to make them discoverable.
Also, just as a heads up, there is a link at the bottom of the search page [0] if anyone would like to build one of these search pages on any topic that might be interesting. It's just a form now bc we are working through the process.
In the broadest sense, a highly targeted search engine (like this) can provide better results because Google has to determine user search intent AND return the right results from trillions of webpages. The advantage of this search engine is that all of the users are looking for the same type of information, and the result sources can be curated to ensure high quality and relevant results.
The more targeted the topic, the harder the time Google has to provide quality signal through the noise of SEO and sheer volume of content on the web.
In addition to the above, the UI provides some cool features like allowing horizontal scrolling of sources to provider higher information density (important for discovery), and some source content can be viewed in the side pane without leaving the page.
But ultimately it would be good to hear if this approach does make it easier to find relevant and higher quality Pytorch info.
Worth mentioning is the Alexandria.org project [0]. It is a non-profit search engine built on data from Common Crawl. The coverage is limited because of Common Crawl, but the relevance is decent. They also provide an API.
I believe one of the biggest impacts toward breaking up Google's monopoly on search is making them open up access to their index, even requiring Google to provide direct API search access for others to build alternative search products. They have a search API today, but it is prohibitively expensive to build on ($5/1000 calls).
I built a fairly popular search engine a couple years back, but the cost of Google's search API and increasing number of bot attacks make it difficult to reason keeping it online.
I just finished reading the PaLM [0] blog post and being genuinely impressed with all the NLP advancements. This search result was a fun juxtaposition.
Mojeek is excellent, and because they use 100% their own index they have a much higher hill to climb.
When I say "Best"for a general search engine, my definition is that it would fulfill the needs of myself and my non-technical family members. Kagi and Brave Search both do that while being different enough to not be just another Bing clone. I use Mojeek often, think it is great, and having their own index is a tremendous asset, but it doesn't quite meet that full definition yet.
Credit to them for trying some new things on the UI front, but it looks like the organic results are from Bing (like most other alt/privacy search engines). It would be interesting to learn more about how/if they plan to build their own index, or set themselves apart.
IMO, Kagi and Brave search are the two best alternative general search engines right now.
Thought that it was interesting that this is a centralized platform, especially considering the coalition of companies. This type of use case has been a typical example of something that a blockchain based solution could be applicable.
This is really an analysis of the use of biased language in news articles, which is interesting but only one dimension of potential bias.
It is very possible to use non "charged" language, but still report a topic with a strong bias. For example, Slate is left leaning by most measures, but the below landscape chart from the study has them dead center. Maybe they are better at using neutral terms?
They key is to only submit information that they already have, not anything new. For documents that need to be uploaded in some cases, the options are either to use a heavily redacted real document with everything blacked out except for the essential info, or just upload a random file because most of the time no one checks and it is just a required field to submit a form.
You are correct it is not actually "deleted", but it will stop your information from showing up on the website.
As an exercise I once went through the process of manually requesting my information removed from most of the top brokers.
It would be difficult to automate because the opt-out processes usually aren't straightforward like unsubscribing from an email list. Many sites make it purposely difficult and involve going through multiple steps, providing verification like a drivers license, and email confirmations.
Nice! Maybe at one point you can release a general web search engine for the Common Crawl corpus? It seems even simpler than this proof of concept, but potentially more useful for people looking for a true full text web search.
There isn't an easy way today to explore or search what is contained in the Common Crawl index.
Thank you for all the kind words! I'm the creator of Runnaroo.
Runnaroo started as just a fun experiment, but it quickly became apparent that you could launch a meta-search engine better than just about everything out there (including DDG [0]), and I was frankly surprised how quickly it was embraced by such a large number of people (it's Show:HN reached the #2 spot [1]).
The challenge then became how to fund the cost of the site in an ethical way in line with the site's core principles. I started looking at different solutions [2], including becoming the first search engine to implement Web Monetization, but I never really even came close.
I believe Runnaroo will live on in some iteration, I just have to figure out what that would be.
Also, if you used Runnaroo and liked it, please don't hesitate to reach out (anything AT runnaroo.com). It has been a solo project, but I'm sure the future will involve more collaboration.
I would also be interested in sharing the story of how Runnaroo evolved over the last year, and the different experiences of launching a search engine if anyone is interested or has a platform for those conversations.
Thank you so much for the kind words! It was very much a labor of love as there was no tracking, ads, or really any significant attempt at monetization.
I was actually happy keeping it up for all of the users such as yourself, but it started to become not worth the effort as it grew and more and more people would abuse the service.
Runnaroo is a one-man side project operation, and I hope it has shown what is possible in the search engine space with minimal resources.
I could try to find some sources, but my guess at the best Web3 sites would probably miss the mark.