FrankChen021 opened a new issue, #20385:
URL: https://github.com/apache/druid/issues/20385

   _This issue was generated automatically by Claude Code (Anthropic's AI 
coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. 
Analysis and suggested fixes are AI-produced; please verify before acting on 
them._
   
   This triage covers the 3 commits merged to master on 2026-09-18 (all 
Dependabot version bumps merged within 3 minutes of each other). One commit 
(3cbd00d9d7, #20363, the last of the three) had failed jobs, two in total: the 
push-triggered `docker-tests` job and the `security vulnerabilities` job of the 
scheduled Cron Job ITs run. Neither failure is caused by the commit, which is a 
one-line `json-flattener` version bump in `benchmarks/pom.xml`. 
`KubernetesClusterDockerTest` timed out waiting for the k3s `druid-coordinator` 
pod to become Ready; the same test passed in the two `docker-tests` jobs that 
ran concurrently for the commits merged 1 and 3 minutes earlier (b0b17848ac, 
3248d73644), on the PR's own pre-merge run (5da3638dd6) and on every other 
master push since 2026-09-10, when the same setup failure last hit (router pod, 
#20312). The security job repeats yesterday's pattern: the Sonatype OSS Index 
returned HTTP 402 on `druid-processing`, and the workflow's fallback then launc
 hed a second, credential-less scan that flagged the known 
`elasticache-java-cluster-client` false positive on `druid-server`. That false 
positive has been present in every scheduled master run since at least 
2026-09-15 (the job itself has been red every day since 2026-09-10) and is a 
persistent CI problem, not a regression from this commit (fix in #20236). No 
re-runs had been triggered; `embedded-tests` run with 
`surefire.rerunFailingTestsCount=0`, so the docker test failed on its single 
attempt.
   
   ### Summary
   
   | Commit | Failed job | Failure log | Root cause | Verdict |
   |---|---|---|---|---|
   | 3cbd00d9d7 (#20363, build(deps): bump json-flattener) | `docker-tests` | 
[job 
105449382571](https://github.com/apache/druid/actions/runs/35296288908/job/105449382571)
 | `KubernetesClusterDockerTest` setup: `ISE: Timed out waiting for 
pod[druid-coordinator-7fdb8bb578-l2jwd] to be ready` after the 300 s 
`POD_READY_TIMEOUT_SECONDS` in `K3sClusterResource.waitUntilPodIsReady`, 
although the coordinator JVM had announced itself in ZooKeeper 21 s after the 
manifests were applied; single attempt (no surefire retries in 
`embedded-tests`); the other 101 docker tests in the job passed | Flaky / infra 
|
   | 3cbd00d9d7 (#20363, build(deps): bump json-flattener) | `security 
vulnerabilities (cron)` | [job 
105467974682](https://github.com/apache/druid/actions/runs/35302536227/job/105467974682)
 | `dependency-check-maven:13.0.0:check`: on `druid-processing`, 
`AnalysisException: Sonatype OSS Index / Guide credits insufficient / payment 
required` (HTTP 402 from `api.guide.sonatype.com/api/v3/component-report`); the 
fallback's unescaped backticks then ran a second scan in which 
`elasticache-java-cluster-client-1.2.4.jar` matched 
`cpe:2.3:a:memcached:memcached:1.2.4` and was flagged with 10 memcached server 
CVEs >= 7.0 (CVE-2016-8704 9.8, CVE-2023-46853 9.8, CVE-2026-47783 8.1, 
CVE-2026-47784 8.1, ...) | Infra (OSS Index 402); Persistent (fix in #20236) |
   
   ### Analysis and suggested fixes
   
   **1. `KubernetesClusterDockerTest` (embedded-tests, `docker-tests`): 
coordinator pod not Ready within 300 s**
   
   `KubernetesClusterDockerTest` extends `IngestionSmokeTest` and runs the 
Coordinator, Overlord, Historical, MiddleManager, Router and Broker as pods 
inside a `rancher/k3s:v1.35.0-k3s1` Testcontainer, each with `DRUID_XMX=128m` 
and `hostNetwork: true`, next to an embedded Overlord, Broker and event 
collector. In `K3sClusterResource.onStarted` the test imports the locally built 
`apache/druid:docker-tests` image into k3s' containerd (01:48:03 to 01:48:39, 
37 s), applies the six Deployments (01:48:40.6 to 01:48:41.3) and then calls 
`waitUntilPodsAreReady`, which iterates the pods in the `druid` namespace and 
waits up to 300 s for each pod's `Ready` condition. The coordinator pod was the 
first one checked and never reported `Ready=True`, so the setup failed at 
01:53:42 and the whole class errored before any test method ran. The failsafe 
output shows that the coordinator, Overlord and Broker JVMs were alive and had 
announced themselves in ZooKeeper at 01:49:02 to 01:49:04 (`Node [http://
 10.1.0.21:8081] of role [coordinator] detected`), i.e. 21 s after the 
manifests were applied, so the image import and pod scheduling worked and the 
Druid process reached the ANNOUNCEMENTS lifecycle stage. The Deployment 
template has no readiness probe, so `Ready` only depends on the kubelet 
reporting the container as running; a container that later exits and restarts 
(a 128 MB heap is tight for a coordinator that loads `druid-s3-extensions`, 
`druid-kafka-indexing-service` and `postgresql-metadata-storage`), or a k3s 
node that flaps to `NotReady` on an overloaded runner, both leave the pod at 
`Ready=False` for long stretches. The code catches only 
`KubernetesClientTimeoutException` and rethrows an ISE with the pod name, so 
neither the pod's `status.conditions`, `containerStatuses` (restart count, last 
termination reason) nor the container log are available, and the 
`failure-druid-container-logs` artifact only contains logs of 
Testcontainers-based tests, not of the k3s pods. This is t
 he same failure mode as the 2026-09-10 router-pod occurrence triaged in #20312.
   
   The commit (#20363) only bumps `com.github.wnameless.json:json-flattener` in 
`benchmarks/pom.xml` and does not touch `embedded-tests`, the Docker image or 
Kubernetes code. The same test passed in the `docker-tests` jobs of b0b17848ac 
and 3248d73644, which ran concurrently on other runners (230 s and 255 s for 
the whole class), on the PR's pre-merge run and on the three master pushes 
since (a007e33a76, 45a4440d24, bf692854a3).
   
   Suggested fix: make the timeout diagnosable and less environment-sensitive. 
In `K3sClusterResource.waitUntilPodIsReady`, on 
`KubernetesClientTimeoutException` include `pod.get().getStatus()` (phase, 
conditions, `containerStatuses[].restartCount` and `lastState.terminated`), the 
pod's events and the tail of `pod.getLog()` in the ISE message, and dump every 
pod's log into the `druid-container-logs` folder in `stop()` so the artifact 
covers k8s tests too. Add an HTTP `readinessProbe` on `/status/health` (with a 
`startupProbe` or `initialDelaySeconds` of about 30 s) to 
`manifests/druid-service.yaml` so `Ready` reflects the Druid process rather 
than only the container state, and consider raising `DRUID_XMX` for the 
coordinator and overlord pods or lowering `druid.coordinator.period` so the 128 
MB heap does not cause restarts. Finally, since `waitUntilServiceIsHealthy` 
already polls `/status/health`, the pod-ready wait could be relaxed to "not in 
`ImagePullBackOff`/`CrashLoopBackOff`" a
 nd let the health poll decide, so a transient `NotReady` blip on the k3s node 
does not fail the class.
   
   **2. `security vulnerabilities (cron)`: Sonatype OSS Index HTTP 402, and 
false-positive memcached CVEs on `elasticache-java-cluster-client-1.2.4.jar`**
   
   The job runs `mvn dependency-check:purge dependency-check:check`; this time 
the NVD update finished in 11 min (394,957 records) and the first module 
scanned, `druid-processing`, failed because the OSS Index analyzer got `402 
Payment Required` from `api.guide.sonatype.com` ("credits insufficient / 
payment required, disabling the analyzer"), which dependency-check 13 treats as 
a fatal analysis exception. The same 402 hit every `security vulnerabilities` 
job on 2026-09-17 and 2026-09-19 (including the `38.0.0-fix-cves` PR runs), so 
it is a quota problem on the OSS Index account, unrelated to the commit. 
Because the fallback in `.github/workflows/cron-job-its.yml` embeds `` `mvn 
dependency-check:check` `` in unescaped backticks inside a double-quoted `echo` 
string, bash executed a second full `dependency-check:check` (without the OSS 
Index and NVD credentials) as a command substitution; that second scan produced 
the `druid-server` failure in which `elasticache-java-cluster-client-1.2.
 4.jar` is matched to `cpe:2.3:a:memcached:memcached:1.2.4` (the C memcached 
server daemon) purely on the coinciding version string and flagged with 10 
server-side CVEs. Druid only uses this jar as a client library in the memcached 
cache. `owasp-dependency-check-suppressions.xml` on master still has no entry 
for it (the `38.0.0` branch got one via #20335), and this false positive has 
appeared in every scheduled master run since at least 2026-09-15 (see also 
#20360, #20366, #20378). Note that today's scheduled run (bf692854a3) 
additionally flags `zookeeper-3.8.6.jar` with CVE-2026-79993, CVE-2026-59739 
and CVE-2026-59969 (all 7.5), so the job will keep failing after the memcached 
suppression lands unless ZooKeeper is bumped or those are triaged too.
   
   Suggested fix: merge #20236, which adds a `<suppress>` entry matching 
`^pkg:maven/com\.amazonaws/elasticache-java-cluster-client@.*$` against 
`cpe:/a:memcached:memcached`, or split that block into a standalone PR so the 
cron job can go green. For the OSS Index outage, replenish or migrate the 
`OSS_INDEX_USERNAME`/`OSS_INDEX_PASSWORD` secrets to a Sonatype Guide personal 
access token (the log warns legacy tokens stop working on 2026-12-31) and pass 
`-DossIndexWarnOnlyOnRemoteErrors=true` so a remote OSS Index error degrades to 
a warning. In the workflow, quote the fallback message with single quotes or 
escape the backticks so it stops launching a second scan. #20126 (cache the NVD 
database between runs instead of purging it) remains worthwhile to keep the 
job's runtime predictable.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to