Am 24.02.2016 um 20:17 schrieb Francisco Andrés Fernández:
Hi all, I'm extracting some text from pdf, through Tika in Solr. As result, some important words end with spaces between characters. For example, I could have the word "Subtitle" that I want to detect, written like "S u b t i t l e".
You could try to modify spacingTolerance or averageCharTolerance in PDFTextStripper (find out if TIKA supports this), but it is likely that if spaces are ignored, they would be ignored at other places where you don't want it.
If possible, please upload your file somewhere. Tilman
How could I make PdfBox detect this type of word occurrence? Many thanks, Francisco
--------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]

