taepper opened a new pull request, #51472: URL: https://github.com/apache/arrow/pull/51472
### Rationale for this change Resolves #51470: `simdjson` has a multi-stage parser that first finds structural elements in the entire buffer (including document boundaries). Only the second stage, which is lazy, actually need to parse the documents. Similarily `parse_many` already computes all document boundaries and we can iterate its output to consume the documents without parsing any documents: https://github.com/simdjson/simdjson/blob/6913ee1a17a57d23a2e9f458d8e04e96cd43ecab/doc/parse_many.md?plain=1#L85-L88 ### What changes are included in this PR? This delimits JSON documents without fully parsing them ### Are these changes tested? Yes. before: ``` ChunkJSONPrettyPrinted 528656 ns 528396 ns 1340 bytes_per_second=394.823Mi/s json_size=218.757k ChunkJSONPrettyPrintedMultipleBlocks 664758 ns 625186 ns 1143 block_size=27.344k bytes_per_second=333.697Mi/s json_size=218.757k ``` after: ``` ChunkJSONPrettyPrinted 91815 ns 91811 ns 7358 bytes_per_second=2.21906Gi/s json_size=218.757k ChunkJSONPrettyPrintedMultipleBlocks 197358 ns 167654 ns 3719 block_size=27.344k bytes_per_second=1.2152Gi/s json_size=218.757k ``` ### Are there any user-facing changes? No. ### Was AI used for this PR? In accordance to the [AI generation guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code), please disclose below whether and how AI was used in this PR. **PR code and description written by:** - [X] Human - [X] AI (assisted in code drafting) **Reviewed before submission by:** - [X] Human - [ ] AI - [ ] Not reviewed -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
