shahar1 commented on code in PR #69855:
URL: https://github.com/apache/airflow/pull/69855#discussion_r4048219991


##########
providers/google/src/airflow/providers/google/cloud/operators/dataproc.py:
##########
@@ -2536,7 +2545,7 @@ def execute(self, context: Context):
                 region=self.region,
                 project_id=self.project_id,
                 batch=self.batch,
-                batch_id=self.batch_id,
+                batch_id=requested_batch_id,

Review Comment:
   A fresh `batch_id` does not ensure a fresh batch when `request_id` is also 
supplied. Each task attempt still sends the same `self.request_id`, and 
[Dataproc 
documents](https://docs.cloud.google.com/managed-spark/docs/reference/rest/v1/projects.locations.batches/create)
 that duplicate request IDs return the original operation. The code below then 
reads the old batch ID from that operation, so a retry can return to the 
previous failed batch without submitting new work.
   Could we reject batch_id_prefix combined with request_id, or generate a 
request ID per task attempt while keeping it stable across network retries? 
Please also cover this interaction in a retry test.
   
   ---
   
   Drafted by Codex with my supervision



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to