We at CloudScrape.com are happy to welcome you onboard our platform - you get 20 hours of scraping runtime for free and the ability to navigate much more complex websites, opening up tons more opportunities then you've had previously. And we're here to stay :)
Kind regards,
Henrik Hofmeister, CTO, Co-Founder @ CloudScrape.com
It's great the see discussions going on here - would like to tie a few comments to the questions of ethical aspects of web scraping:
As some has pointed out scraping is not exactly a new thing and a lot of the biggest sites out there are built on the basis of web scraping or crawling. We provide a tool and expect you use that tool while abiding the law - and if not we will of course shut your account down immediately. Breaking the law includes violating copyrights and performing DDoS attacks (Although they will be rather small attacks since even 50 concurrent agents is no big deal for most websites).
We consider ourselves good netizens. We wish nothing more than to provide a good, easily accessible and safe tool for extracting valuable information from the internet, be it for a price comparison site in a market that lacks transparency, business intelligence for your company to make informed and wiser decisions, or a PhD project that requires access to millions of data points available online in unstructured form.
Additionally if you feel we're providing services that has ill-intent - we are not providing any services (Captcha and proxy rotation) that anyone with a bit of programming skill can not easily use in their own software. The main difference is that we are actively improving and focusing not only on making a good experience for our users - but also on minimizing the impact on the sites being scraped. This involves several things like automated throttling and slow-site detection, request caching, and blocking requests to services such as google analytics - to not interfere with site owners stats.
Thank you for the recommandations! Would love to know more about your exact use cases - we're aiming for the 100% :) If you'd reach out on [email protected] - we'd be happy to investigate if we can't make CloudScrape work for you as well.
Thank your for the kind words :) While you're right that we do run a headless browser-ish thing what sets us apart from most of our competitors, other than the point-and-click approach is that we autodetect everything that's going on in the browser - which also means that in most cases you dont need to know what's going on. What this effectively means is that you'll often spend no time reverse-engineering and be able to scrape even wildly complex javascript-heavy sites in minutes instead of hours - and have them be a lot more stable than they would otherwise since there's no "Wait for 5 seconds" which will only be enough 95% of the time.
We see a lot of our clients manage to do what they need done - even daily scrapes - for as little as $29 / month since scraping a news site daily will often take up no more than a few minutes.
We've just launched our SaaS web scraping product called CloudScrape, which will let you scrape any page - no matter the complexity, using our in-browser editor and webkit-powered runtime. Try it out for free! Feedback / comments are very welcome and appreciated.
We'll be adding more languages and frameworks as soon as we can!