shivaam commented on PR #70455:
URL: https://github.com/apache/airflow/pull/70455#issuecomment-5956216227

   I’m closing this for now. Batching improves publishing performance, but for 
the low-latency Redis setup we tested, I don’t think the savings justify the 
added failure-handling and maintenance complexity. I’ll keep the prototype and 
results in case we revisit it.
   
   The publishing tests showed these results:
   
   | Messages published together | Condition | Without batching | With batching 
|
   |---:|---|---:|---:|
   | 16 | Direct Redis connection | 4.25 ms | 2.29 ms |
   | 1,000 | Direct | 136.85 ms | 67.63 ms |
   | 10,000 | Direct stress test | 1,072.09 ms | 548.03 ms |
   | 1,000 | Added 1 ms response delay | 370.74 ms | 92.26 ms |
   | 1,000 | Added 5 ms response delay | 980.90 ms | 143.46 ms |
   
   These are median publishing times, not task execution times. The comparisons 
used selected process configurations rather than the default settings on both 
sides.
   
   Batching helped, particularly with larger groups and slower Redis responses. 
However, we expect network latency to be low when Redis runs close to the 
scheduler. Different Pods may run on the same machine or within the same 
availability zone. Our separate-container setup on one VM measured a median 
Redis ping of **0.058 ms**. AWS describes same-zone network round trips as 
typically **below 1 ms**, although that is not a guarantee for Redis response 
times. [AWS 
guidance](https://aws.amazon.com/blogs/architecture/improving-performance-and-reducing-cost-using-availability-zone-affinity/).
   
   The 1 ms and 5 ms delays were deliberately added to test slower responses; 
we should not assume every deployment experiences them. Without those delays, 
batching saved about **2 ms for 16 messages** and **69 ms for 1,000 messages**.
   
   There is also additional failure-handling complexity. Our tests showed that 
Redis can accept some messages even when the pipeline raises an exception. 
Retrying the whole batch can duplicate delivery, while treating everything as 
failed can incorrectly report accepted messages as failures.
   
   I’d revisit this if a deployment with consistently slower Redis responses 
shows a clear need for it. Thanks for the review and feedback.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to