Hi all,

Flume 2.0.0 is getting close, and before I set up the release process
I'd appreciate your advice on how to structure it. Many of you have far
more experience than me with releasing multi-repository projects like
Log4j, so feel free to point out anything that looks off.

Flume is now organized like Log4j:

- `logging-flume`: the core components, including the executable Flume
Node. It currently also hosts `flume-parent`, the BOMs and the
(disabled) binary distribution.
- 14 satellite repositories with the sources, sinks and channels that
carry heavy dependencies (Avro, Thrift, Hadoop, Kafka, Spring Boot, ...).

I expect the core to change rarely, while the satellites will need
releases that follow the lifecycle of their dependencies. So I'd like to
release the satellites independently, without waiting for (or forcing) a
core release.

The catch is that the core repository still points at the satellites:
its BOM lists their artifacts, and the distribution bundles them, so a
complete, consistent core release would have to come after every round
of satellite releases.

## What I'm planning

My idea is to make dependencies point one way only: satellites depend on
the core, never the reverse. Everything that aggregates the satellites
would move to a new repository, released last.

1. `logging-flume` -> **Apache Flume Core** (ATR key `logging-flume-core`)
   - Core modules only. The non-trivial parts of `flume-parent`
(Surefire, PMD, Spotless configuration) would move to `logging-parent`,
and the dependencies only some satellites need for tests (e.g., the
Hadoop mini-cluster) would move to those satellites. What remains is
little more than metadata, so `flume-parent` can be dropped and all
repositories can inherit from `logging-parent` directly.
   - A `flume-core-bom` listing only the artifacts this repository builds.
   - No more references to satellite artifacts.
2. Satellite repositories -> one subproject each. They depend on a
released core through `flume-core-bom` and have their own release cadence.
3. New `logging-flume-dist` -> **Apache Flume**, released whenever there
is a tested set worth publishing:
   - `flume-bom`, listing compatible versions of the core and all
satellites;
   - a binary distribution recipe that users can fork and extend with
their own plugins;
   - a Docker image definition.

None of the BOMs has been released yet, so they can still be renamed or
split.

The release order for 2.0.0 would be: core, then satellites, then dist.

## Binary distribution and Docker

The 1.x ZIP bundled every module, yet users usually had to add their own
plugins anyway. So I'm leaning towards publishing a recipe (an assembly
project to copy) and a Docker image instead of a monolithic ZIP.

For Docker, my reading of the release policy is that convenience
binaries must only contain artifacts built from the released sources.
Rebuilding the image on an updated base image doesn't change any Flume
bits, so I think it can be rebuilt on a schedule from CI without a new
vote, with a vote only when the recipe or the Flume version changes. I'm
not sure about this, though.

## Where I'd like your advice

1. Does this split make sense to you, or would you organize it
differently? In particular, are you fine with moving the Flume build
configuration to `logging-parent`?
2. Should the satellites keep the core version number (2.0.x) or get
independent version numbers? I would rather version them separately.
3. Is a binary ZIP still worth producing, or are a recipe and a Docker
image enough?

Thanks!

Piotr

Reply via email to