taepper opened a new pull request, #51473: URL: https://github.com/apache/arrow/pull/51473
### Rationale for this change Resolves #51471: When invoking `simdjson`'s `parse_many`, it uses threading to perform its stage 1 parsing on the next batch when while parsing the current batch: https://github.com/simdjson/simdjson/blob/master/doc/parse_many.md?plain=1#L96-L104 In most cases, we do not need to parse any batches after the first one to find delimiters (in fact, we currently would even error, when the first batch does not contain the whole first document). By disabling the parser's threading we can expect performance improvements ### What changes are included in this PR? This disables threading in the parser by setting `parser_.threaded = false;` ### Are these changes tested? Yes before: ``` ChunkJSONPrettyPrintedMultipleBlocks 687263 ns 648184 ns 1099 block_size=27.344k bytes_per_second=321.858Mi/s json_size=218.757k ``` after: ``` ChunkJSONPrettyPrintedMultipleBlocks 561073 ns 561058 ns 1309 block_size=27.344k bytes_per_second=371.839Mi/s json_size=218.757k ``` only #51470: ``` ChunkJSONPrettyPrintedMultipleBlocks 197358 ns 167654 ns 3719 block_size=27.344k bytes_per_second=1.2152Gi/s json_size=218.757k ``` together with #51470: ``` ChunkJSONPrettyPrintedMultipleBlocks 108151 ns 108148 ns 6463 block_size=27.344k bytes_per_second=1.88383Gi/s json_size=218.757k ``` ### Are there any user-facing changes? No ### Was AI used for this PR? **PR code and description written by:** - [x] Human - [ ] AI **Reviewed before submission by:** - [x] Human - [ ] AI - [ ] Not reviewed -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
