I'd go with .spread(...), as both batch(...) and shard(...) are already semantically overloaded.
Shahar On Thu, Sep 17, 2026 at 12:26 PM Blain David <[email protected]> wrote: > Hi all, > > While reviewing the AIP-104 PR ( > https://github.com/apache/airflow/pull/62922), Amogh raised a naming > concern about the Dynamic Task Batching API that I would like to settle > here before the PR merges, > since renaming afterwards would be a breaking change. > > ## What `.batch(size=N)` does today > > `task.batch(size=17).iterate(url=urls)` creates 17 mapped task instances > through Dynamic Task > Mapping, and each of them iterates over its share of `urls` in a single > task instance using > Iterable Tasks. The items are distributed round robin: item `i` goes to > task instance `i % 17`. > So `size` is the number of task instances, not a chunk length, and the > task instances do not > receive contiguous slices of the input. > > The round robin scheme is deliberate. The number of task instances has to > be fixed before the > underlying iterable is consumed, because the scheduler needs the mapping > cardinality up front. > With round robin that number is simply `size`. Contiguous chunking, the > way `itertools.batched` > works, would need `ceil(len(items) / size)` task instances, which is > unknowable until a possibly > unbounded or paginated iterable has been fully drained. That would defeat > the purpose of iterating > lazily. > > ## The concern > > The name reads like `itertools.batched(iterable, size)`, where `size` is > the chunk length and the > number of chunks is derived from it. Ours is the opposite: the number of > task instances is given, > and the per-instance share is derived. The class docstring and the Task > SDK docs currently spend a > full paragraph undoing that first impression. A name that carries the > meaning directly would not > need it. > > ## Options > > 1. Keep `.batch(size=N)` and rely on documentation. Lowest churn, but the > mismatch with > `itertools.batched` stays and every new reader has to be corrected. > > 2. Keep the method, rename the parameter: `.batch(count=N)` or > `.batch(tasks=N)`. Keeps the > "Dynamic Task Batching" vocabulary from the AIP, and `count` or `tasks` > says what the number > is. `.batch(tasks=17).iterate(url=urls)` reads as "spread over 17 > tasks, then iterate". > > 3. Rename the method as well, to something that describes partitioning > rather than chunking, for > example `.partition(count=N)`, `.shard(count=N)` or > `.spread(across=N)`. Most descriptive, but > drifts from the AIP's own terminology, and "shard" and "partition" > carry data-engineering > connotations that may suggest contiguous ranges just as much as "batch" > does. > > My preference is option 2 with `count=`. It keeps the API shape and the > AIP terminology, fixes > the misleading part, and costs nothing now because the feature is > unreleased. `size` would be > kept as a rejected alias raising a clear error rather than silently > accepted, so nobody copies > the old spelling from an early draft. > > Unless there are objections or a better name comes up, I will treat this > as lazy consensus in > 72 hours and update the PR, the docs and the AIP page accordingly. > > Thanks, > David > > General (Internal Property) >
