thx, recognized the variance, two things: results are based on a tech domain specific vector (based on 1.2m tech articles / 20y), so non-techs fall off. We source content on requested entities live. Descriptions of younger entities are more concise/less global -> more valid classification. But then again, our assumption was that there is more demand for classification of little known entities than for F100s. But we are here to validate/falsify that assumption.
we start with some automated content acquisition, before we use nlp for keyword extraction and tagging. The company-tag description is then located in a huge tech-domain-specific vector space. From vector space similarities and a couple of heuristics, we derive the company classification.