(literally just scale it up x2 or x4).
Just to clarify, the input image should be 300dpi. In this case, the use is extracting words that can be used in full-text search, so structural extraction isn't a key criteria.
In case someone wants to know more, the former is known as "full page OCR" and the latter as "data capture"/"document processing" (or IDP, intelligent document processing).
Is it because you want to own the IP or for learning purposes?