Hi Team,

Kindly please look into the below issue and it would be great if you guys
give a solution for this

Thanks,
Saradhi

---------- Forwarded message ---------
From: vijaya saradhi reddy <[email protected]>
Date: Mon, Jul 6, 2020 at 8:25 PM
Subject: Unstructured Extraction by tika(Pdf)
To: <[email protected]>


Hi Chris,
Please help me out from this troublesome issue
For Pdf data extraction i found tika as the best library comparing to pdf
plumber and PyMUPdf(fitz), but i am facing a small issue while trying to
extract data from below pdf
[image: image.png]


>From the above pdf Apache tika extracting data like below image

[image: image.png]


Its extracting as above but i want my output as below image as it should
extract as it is like in pdf. Below extraction results are using Pdf
plumber, can i get the below result using apache tika. Please help me out
from this as iam spending lots of time on tis

[image: image.png]

Reply via email to