CurtHagenlocher commented on issue #409:
URL: https://github.com/apache/arrow-dotnet/issues/409#issuecomment-5851238527

   Thanks for the detailed report. I don't think it's Arrow that's writing the 
zeros. The pattern better fits something outside your process shrinking the 
file back to an earlier size while it's still being written.
   
   **Can ArrowFileWriter advance Position without writing bytes?**
   
   No. ArrowFileWriter and ArrowStreamWriter never call Seek or SetLength on 
the base stream. They only read
   BaseStream.Position, to record the block offsets for the footer. All padding 
is written as real zero bytes, not by
   seeking past it. So Position only moves when a Write call succeeds.
   
   The continuation marker (0xFFFFFFFF) and the message header come from 
Arrow's own memory, filled just before each
   write. That memory isn't released or reused until the write returns. This 
also holds on net462/net472, where writes go
   through a temporary pooled copy. Your custom MemoryManager<byte> could at 
worst affect the data values in the body.
   It can't zero out the marker or header. Your Rx usage also looks fine: 
Buffer(timeSpan, count) hands off each batch
   under a lock, so WriteRecordBatch calls don't overlap even when they arrive 
on different threads.
   
   A few details from open-ephys/bonsai-onix1#683 stand out:
   
   - The zeroed ranges start and end exactly on message boundaries (for example 
0x323C50 to 0x43DB00). Those offsets
     aren't aligned to disk sectors or pages. A disk or cache fault loses whole 
sectors or pages, so it wouldn't line up
     with Arrow's messages in file after file.
   - Every zeroed range is one or more whole batches, always early in the 
recording.
   - The file with no footer holds exactly four complete batches, with no 
partial one.
   - The original machine was saving to a Synology Drive synced folder.
   
   One explanation covers both failures: another process shrinks the file back 
to an earlier size while your writer still
   has it open. Suppose the file is cut back to the end of batch 2 after batch 
3 has been written. Your FileStream still
   points past batch 3, so batch 4 lands at the correct offset and Windows 
fills the gap with zeros. That's exactly what
   you're seeing. The same thing happening after the footer is written leaves a 
file of whole batches with no footer.
   
   The writer is idle between batches, so the size the file gets cut back to is 
almost always a batch boundary. A sync
   client that notices a new, growing file and acts on it soon after it appears 
would explain why this only happens
   early. Your FileStream is opened with FileShare.Read, so an ordinary program 
couldn't open the file for writing. That
   points at something with a kernel-mode driver, such as a sync client's 
on-demand sync driver or antivirus software.
   
   I haven't reproduced this, so please treat it as a hypothesis. It would 
explain why only one machine is affected.
   
   **How to confirm it**
   
   On the affected machine, run Process Monitor with a filter on the output 
folder or file name. Look for
   `SetEndOfFileInformationFile`, `SetAllocationInformationFile` or `WriteFile` 
on the `.arrow` file from any process other than
   Bonsai, especially the Synology Drive processes. Synology Drive's client 
logs from around the start of the recording
   may also help.
   
   **Would FileOptions.WriteThrough or Flush(true) help?**
   
   Probably not. Those settings protect against losing data in a crash or power 
loss. They can't stop another process
   from truncating or rewriting the file. The practical fix is to record to a 
folder that isn't synced and move each file
   into the synced folder once it's closed. Alternatively, write under a 
temporary name or extension the sync client is
   set to ignore and rename the file when it's complete.
   
   If Process Monitor shows the file being changed by something else, please 
post what you find here. If it shows nothing
   unusual, we can try to dig a little further on the Arrow side.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to