Hello, I have had some reasonable success with 'pdfquery' if you like Python. It works with regional text as well. Also, for tabular data, do try pdf-table-extract if quick and dirty works for you.
Java folks should try pdfbox. On 20 January 2017 at 15:23, mohit ranjan <[email protected]> wrote: > Sorry if this is off-topic, but have seen threads here about liberating > data from PDFs. > Most likely there will be lot of scanned PDFs among them. > > Do we have any in-house expert on this and which library/tool (preferably > not paid) to extract tables in scanned PDF/JPG ? > > CVision > <http://www.cvisiontech.com/library/ocr/file-ocr/ocr-table-recognition.html> > does a decent job, but it's paid. > > > > - Mohit > > -- > Datameet is a community of Data Science enthusiasts in India. Know more > about us by visiting http://datameet.org > --- > You received this message because you are subscribed to the Google Groups > "datameet" group. > To unsubscribe from this group and stop receiving emails from it, send an > email to [email protected]. > For more options, visit https://groups.google.com/d/optout. > -- Datameet is a community of Data Science enthusiasts in India. Know more about us by visiting http://datameet.org --- You received this message because you are subscribed to the Google Groups "datameet" group. To unsubscribe from this group and stop receiving emails from it, send an email to [email protected]. For more options, visit https://groups.google.com/d/optout.
