I'm sure people are more likely to link to the actual SICP page [1] rather than the Amazon page, especially as it's free on the website. If the discussion was in the context of buying the book, I would personally still link the MIT Press page [2], rather than the Amazon page.
> I quickly discovered that such a system was overkill for me, and resorted to using an open source implementation of a simpler algorithm [1].
Maybe worth pointing out that the "simpler algorithm" they used seems to be a cascading ensemble of adaptive boosting algorithms, technique similar to the ones used on Kaggle to win the big prizes. Maybe simpler than some neural nets in some ways, but nothing close to the simplicity of nearest neighbour search.
Both arithmetic mean and median are averages, so unless they specify what average they used for the annual income, you may only assume what the author meant.
EDIT: Never mind, the author also says:
>Moreover, we compare median annual values for software engineers with mean annual values for the general population
Live attack To obtain an exact measurement of our attack’s accuracy, we run our automated captcha-breaker against reCaptcha. We employ the Clarifai service as it shows the best result amount other services.
Labelled dataset. We created a labelled dataset to exploit the image repetition. We manually labelled 3,000 images collected from challenges, and assigned each image a tag describing the content. We selected the appropriate tags from our hint list. We used pHash for the comparison, as it is very efficient, and allows our system to compare all the images from a challenge to our dataset in 3.3 seconds. We ran our captcha-breaking system against 2,235 captchas, and obtained a 70.78% accuracy. The higher accuracy compared to the simulated experiments is, at least partially, attributed to the image repetition; the history module located 1,515 sample images and 385 candidate images in our labelled dataset.
Average run time. Our attack is very efficient, with an average duration of 19.2 seconds per challenge. The most time consuming phase is running GRIS, consuming phase, as it searches for all the images in Google and processes the results, including the extraction of links that point to higher resolution versions of the images.
How come? I know What.CD closed down, but I remember that RuTracker had quite a large selection of high-fidelity music records.