I will definitely check out your links. Out indexing Google or Microsoft live search is definitely out of the question for me just because of the amount of data among others. Social search with like Yahoo BOSS might work. I will look into fact-based search using social search and conventional search or some kind of combo.
Good reply. Thanks for replying. I personally think it ultimatily boils down to content that the search engine has. That is the index in this case for Google. I think the semantics and the semantic web will make a huge difference. How I look at it is that because Google's result for "c# string replace" http://www.google.com/search?hl=en&q=c%23+string+replace is much better than Mahalo's http://www.mahalo.com/C#_string_replace, Google is good. I was thinking about how the search will be in the next 2-3 years. The main problem I see with is the discovery of new content in search engine, and I think Google basically brute forced the whole thing by trying to index everything and hope that the results are there and which works alright for the query above. It basically boils down to 100 thousand crawlers and huge index. If social search engine cannot discover this page http://msdn.microsoft.com/en-us/library/fk49wtc1.aspx for this query "c# string replace" for instance, it is game over. Google is alive because of these kinds of results. My main concern is how to discover new content and without the burden of updating the index and bruce forcing the whole thing by indexing every word on a webpage. I know Google indexes pieces of webpages but still it is ton of words to index. I just see huge problem with creating huge index like Google and maintaining that index which is also a lot of work. Also I don't have the resources (money) to create a huge index, which is another main reason.
I wanted to post a general discussion going about creating a search engine since the Ycombinator crowd is a good crowd. How do you guys see Google scaling in comparison with how the web is expanding so much? Do you think they will keep indexing the web and adding new content? I personally don't see that to be very scalable. What do you think Google is thinking about the future in terms of incorporating so many webpages into its index. Obviously Google is using a "pull" mechanism where it crawls and indexes. Do you guys see like "push" mechanism like RSS that will work better and how would that affect the search result? Is there any way besides index to achieve a different kind of search engine? Another kind of content retrieval system is like the Digg where the user "pushes" the content to Digg and therefore it is supposedly more relevant and interesting, but the disadvantage of that becoming a search engine is there isn't a lot of content that the user submits compared with like Google. For instance Digg cannot support query like "c# string replace" while Google will do that very easily because it crawled and indexed the MSDN api pages already, while Digg users might never submit that same page to Digg. My main concern is supporting so many content with different query and being very comprehensive search engine like Google without this huge index and crawling restriction? Any clues? I'm not dreaming about this and I actually want to make it a reality somehow.
This UI design is critical to search engine is just nonsense. All that matters in search engine is the result. You can have the crappiest interface and the best result, you will become billionaire. Simple as that. Design in search engine doesn't matter at all. All that matters is the result.
I'm not concerned about how the newspapers and producers like MSNBC and Washington Post create the story. The main problem is how do the search engine or whatever can search that news. The problem with Google News is that google is crawling and indexing all the webpage news like it was a webpage and it is not very up to date. If there was a fire right now, you won't see that in google right now. On the other hand, RSS is a different topic. First of all not all news producers have RSS. My main question is, if there is a flooding happening right now, what is the best way to find that information. You can't use google now, because they haven't crawled and indexed it. I'm thinking of news search/aggregator realtime. Another nice thing would be if I search "flooding" it should display many news sources so that I can get some perspective. What is the best way to approach this problem. Crawling and indexing is not the way to go I think if it is needed to be instant.
1. Don't spend any huge money on this project until it picks up significantly.
2. treat this as a hobby, which means don't lose sleep over it.
3. focus on your work that makes money if money is a factor
4. take it easy. leave the company as it is.
5. try to make this not be written in LISP. that is a lot of lines of code. convert it.
6. Keep everything going the same time. don't quit your job.