RaphaelAdewale opened a new pull request, #11838:
URL: https://github.com/apache/seatunnel/pull/11838
### Purpose of this pull request
Implements the **Splunk Sink** connector, the first half of #10753. The
Source
will follow in a separate PR.
The sink writes SeaTunnel rows to a Splunk index through the HTTP Event
Collector. Each row becomes one HEC event envelope — the row under `event`,
with `index` / `source` / `sourcetype` / `host` / `time` populated from sink
options — and envelopes are POSTed to `/services/collector/event` in batches.
Part of #10753
### Does this PR introduce _any_ user-facing change?
Yes — a new `Splunk` sink plugin. No existing option, default, or API is
changed. Docs added at `docs/en/connectors/sink/Splunk.md` and the zh
equivalent.
### Design notes
- **Not built on `connector-http-base`.** Its `HttpClientProvider` always
builds `HttpClients.createDefault()` with no TLS hook, and Splunk HEC
commonly sits behind the self-signed certificate generated at install time.
`SplunkHecClient` adds `tls_verify_certificate` / `tls_verify_hostname`
(named to match the Elasticsearch sink).
- **Retry classification.** Only transport errors, HTTP 429 and 5xx are
retried. A bad token, a forbidden index, or a malformed payload fails the
task immediately instead of burning retries on an error that cannot resolve
itself. The buffer is cleared only after the collector accepts the batch,
so
a failed attempt never silently drops events.
- **No connector-level `flush_interval`.** Periodic flushing uses the
engine-level `sink.flush.interval` via `context.registerFlushAction`,
consistent with the direction taken in connector-prometheus and the
Elasticsearch/Hudi sinks. Zeta only; Spark and Flink flush on
`max_batch_size`, checkpoint and close.
- **`time_field` accepts** `TIMESTAMP` (read as UTC), `TIMESTAMP_TZ`, and
`BIGINT` (epoch millis). Any other type fails at startup with a message
naming the field and its type. The value is sent as epoch seconds with
millisecond precision, written in plain notation — Splunk rejects the
scientific notation Jackson produces for doubles at epoch magnitude.
- **Delivery is at-least-once** and documented as such; HEC offers no
server-side dedup for this endpoint.
### How was this patch tested?
- 34 unit tests: config validation and endpoint normalisation, HEC envelope
serialization (metadata mapping, timestamp conversion, plain-notation wire
format), batching and flush thresholds, retry classification, and the
factory OptionRule.
- E2E `SplunkIT`: `splunk/splunk:9.3.2` container, FakeSource → Splunk sink,
asserting the events land in the target index with the configured metadata.
- Additionally verified by hand against a real Splunk 9.3.2 instance using
the
built distribution: a Zeta job wrote 5 rows, and a Splunk search confirmed
the event bodies, `host` from `host_field`, static `source`/`sourcetype`,
and `_time` from `time_field`. An invalid-token run confirmed the fail-fast
path.
Note: the local Docker version (29.3.0) is incompatible with this repo's
pinned Testcontainers client, so no E2E in the repo can run locally — the
pre-existing `HttpIT` fails identically. `SplunkIT` therefore runs for the
first time in CI; the manual round-trip above covers the same ground.
### Check list
* [x] Code changed are covered with tests
* [x] If any new Jar binary package adding in your PR, please add License
Notice according [New License
Guide](https://github.com/apache/seatunnel/blob/dev/docs/en/contribution/new-license.md)
— no new third-party dependency; `httpclient`/`httpcore` are already in
`known-dependencies.txt`
* [x] If necessary, please update the documentation to describe the new
feature
* [x] Update the `plugin-mapping.properties` and add new connector
information in it
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]