rzo1 commented on PR #2183: URL: https://github.com/apache/stormcrawler/pull/2183#issuecomment-5812728965
@tballison Good questions. In Storm every bolt executor is a single thread and `execute()` handles one tuple at a time, so a `ParserBolt` never has more than one parse in flight. Concurrency comes from the bolt's parallelism, which is why I went with one fork per bolt instance. A second client would just sit there idle. For memory, I'd tell users to always set `parser.tika.pipes.jvmargs`, e.g. `-Xmx512m`, and size the host as worker heap plus one fork heap (and some JVM overhead) per `ParserBolt` instance on that host. In my Docker test, 512m was fine for normal documents and for a 13MB PDF with 2,700 pages. A nasty 12MB single-page PDF ran out of memory and came back as `parse crash`, which is what we want. One thing I noticed while looking into this: if no heap is set, Tika gives each fork `-XX:MaxRAMPercentage=60`, since `numClients` is 1. With a few `ParserBolt` instances on the same host that's way more than 100% of RAM, and Tika's overcommit warning can't catch it because each bolt has its own parser. I think `ParserBolt` should set its own default when the user doesn't, either a plain `-Xmx512m` or a percentage split across the `ParserBolt` tasks in the worker. Any preference? Also worth documenting: `topology.message.timeout.secs` needs some headroom beyond `parser.tika.timeout`. Tuples queue up behind a slow parse, and restarting the fork after a timeout took about 1.5s here. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
