taepper opened a new pull request, #51472:
URL: https://github.com/apache/arrow/pull/51472

   ### Rationale for this change
   
   Resolves #51470:
   
   `simdjson` has a multi-stage parser that first finds structural elements in 
the entire buffer (including document boundaries). Only the second stage, which 
is lazy, actually need to parse the documents.
   
   Similarily `parse_many` already computes all document boundaries and we can 
iterate its output to consume the documents without parsing any documents:
   
https://github.com/simdjson/simdjson/blob/6913ee1a17a57d23a2e9f458d8e04e96cd43ecab/doc/parse_many.md?plain=1#L85-L88
   
   ### What changes are included in this PR?
   
   This delimits JSON documents without fully parsing them
   
   ### Are these changes tested?
   
   Yes.
   
   before:
   
   ```
   ChunkJSONPrettyPrinted                                             528656 ns 
      528396 ns         1340 bytes_per_second=394.823Mi/s json_size=218.757k
   ChunkJSONPrettyPrintedMultipleBlocks                               664758 ns 
      625186 ns         1143 block_size=27.344k bytes_per_second=333.697Mi/s 
json_size=218.757k
   ```
   
   after:
   
   ```
   ChunkJSONPrettyPrinted                                              91815 ns 
       91811 ns         7358 bytes_per_second=2.21906Gi/s json_size=218.757k
   ChunkJSONPrettyPrintedMultipleBlocks                               197358 ns 
      167654 ns         3719 block_size=27.344k bytes_per_second=1.2152Gi/s 
json_size=218.757k
   ```
   
   ### Are there any user-facing changes?
   
   No.
   
   ### Was AI used for this PR?
   
   In accordance to the [AI generation 
guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code),
 please disclose below whether and how AI was used in this PR.
   
   **PR code and description written by:**
   
   - [X] Human
   - [X] AI (assisted in code drafting)
   
   **Reviewed before submission by:**
   
   - [X] Human
   - [ ] AI
   - [ ] Not reviewed
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to