rzo1 commented on PR #2183:
URL: https://github.com/apache/stormcrawler/pull/2183#issuecomment-5812728965

   @tballison Good questions. In Storm every bolt executor is a single thread 
and `execute()` handles one tuple at a time, so a `ParserBolt` never has more 
than one parse in flight. Concurrency comes from the bolt's parallelism, which 
is why I went with one fork per bolt instance. A second client would just sit 
there idle.
   
   For memory, I'd tell users to always set `parser.tika.pipes.jvmargs`, e.g. 
`-Xmx512m`, and size the host as worker heap plus one fork heap (and some JVM 
overhead) per `ParserBolt` instance on that host. In my Docker test, 512m was 
fine for normal documents and for a 13MB PDF with 2,700 pages. A nasty 12MB 
single-page PDF ran out of memory and came back as `parse crash`, which is what 
we want.
   
   One thing I noticed while looking into this: if no heap is set, Tika gives 
each fork `-XX:MaxRAMPercentage=60`, since `numClients` is 1. With a few 
`ParserBolt` instances on the same host that's way more than 100% of RAM, and 
Tika's overcommit warning can't catch it because each bolt has its own parser. 
I think `ParserBolt` should set its own default when the user doesn't, either a 
plain `-Xmx512m` or a percentage split across the `ParserBolt` tasks in the 
worker. Any preference?
   
   Also worth documenting: `topology.message.timeout.secs` needs some headroom 
beyond `parser.tika.timeout`. Tuples queue up behind a slow parse, and 
restarting the fork after a timeout took about 1.5s here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to