SEZ9 opened a new pull request, #11982:
URL: https://github.com/apache/seatunnel/pull/11982

   ### Purpose of this pull request
   
   Phase 1 of #11980.
   
   The REST API only reports a job's *current* status. During incident triage 
the
   first three questions cannot be answered from the engine at all:
   
   1. "The job shows RUNNING — has it silently restarted?"
   2. "Is a pipeline crash-looping?"
   3. "How long did it wait for resources before running?"
   
   Today all three require reading node logs; `JobMetricExports` only publishes 
a
   `job_count` gauge by status, and `PhysicalPlan#reportJobStateEvent` pushes
   terminal states only unless `report-non-terminal-job-state` is enabled.
   
   Both signals needed are **already tracked and simply never surfaced**:
   
   - `SubPlan#prepareRestorePipeline()` increments `pipelineRestoreNum` on every
     pipeline restore and gates further restores on `pipelineMaxRestoreNum`; the
     getter is used only for that gate and a log line.
   - `IMAP_STATE_TIMESTAMPS` holds a `Long[]` of per-state entry timestamps for
     **both** job level (key `jobId`) and pipeline level (key 
`PipelineLocation`).
     REST reads exactly one slot of it today — `SCHEDULED`, to render 
`startTime`.
   
   This PR exposes them in `GET /job-info/{jobId}` under a new `diagnostics` 
field.
   
   Why the pipeline part matters: a pipeline restart does **not** change the job
   status (the job stays `RUNNING` across `SubPlan#reset()`), and at job level a
   `RUNNING → FAILING → RUNNING` loop is impossible because `FAILING` can only
   proceed to the terminal `FAILED`. So job-level data alone cannot answer 
question
   2 — `pipelines[].restoreCount` is what makes a crash loop visible.
   
   Design notes:
   
   - **Pure read.** No new recording and no write on the state-transition path.
     `updateStateTimestamps` and `updatePipelineState` are untouched.
   - The master builds the block from its own physical plan; a request served 
by a
     non-master member fetches it through the new `GetJobDiagnosticsOperation` 
so
     the payload is identical on every node.
   - Diagnostics are auxiliary: any failure (master switch in flight, job 
already
     finished, older master that does not know the new operation id) omits the
     field instead of failing the job-info request.
   - The new operation id is appended (17) and existing ids are untouched.
   
   Bounded state-transition history (Phase 2 of #11980) is intentionally **not**
   part of this PR — it needs a write on the transition path and design 
agreement.
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes, purely additive. `GET /job-info/{jobId}` (and the endpoints sharing
   `convertToJson`, i.e. `/running-jobs` and the deprecated 
`/running-job/{jobId}`,
   on both the v1 and v2 REST planes) gains a `diagnostics` object for running
   jobs:
   
   ```json
   "diagnostics": {
     "jobId": "962858126931165185",
     "generatedAt": 1755000004000,
     "stateTimestamps": {
       "INITIALIZING": 1755000000000,
       "CREATED": 1755000000200,
       "SCHEDULED": 1755000001000,
       "RUNNING": 1755000003000
     },
     "pipelines": [
       {
         "pipelineId": 1,
         "pipelineStatus": "RUNNING",
         "restoreCount": 7,
         "maxRestoreCount": 100,
         "stateTimestamps": {
           "SCHEDULED": 1755000001100,
           "DEPLOYING": 1755000002000,
           "RUNNING": 1755000003500
         }
       }
     ],
     "totalPipelineRestoreCount": 7
   }
   ```
   
   `restoreCount` growing while `jobStatus` stays `RUNNING` is a crash loop;
   `SCHEDULED - CREATED` is the resource wait. No existing field changes, and 
the
   field is omitted when it cannot be read, so clients that ignore it are
   unaffected.
   
   Docs updated: `docs/{en,zh}/engines/zeta/rest-api-v2.md` (response example 
plus a
   field-description table) and a cross-reference note in
   `docs/{en,zh}/engines/zeta/rest-api-v1.md`.
   
   ### How was this patch tested?
   
   - New unit test `JobRuntimeDiagnosticsTest` covering: both job and pipeline
     signals rendered; states never entered omitted; a job with no `JobMaster`
     still reporting its state timestamps; a failing pipeline lookup isolated so
     the payload is still produced; cleaned-up timestamps rendering as an empty
     object; and a timestamps array shorter than the current enum (rolling 
upgrade)
     not overflowing.
   - New E2E assertion `RestApiIT#testGetJobDiagnosticsOfRunningJob` runs 
against
     both cluster members, so it covers the master (local build) and the 
non-master
     (master operation) path and asserts they return the same payload.
   
   ### Check list
   
   * [x] If any new Jar binary package adding in your PR, please add License 
Notice according
     [New License 
Guide](https://github.com/apache/seatunnel/blob/dev/docs/en/developer/new-license.md)
   * [x] If necessary, please update the documentation to describe the new 
feature. https://github.com/apache/seatunnel/tree/dev/docs
   * [x] If necessary, please update `incompatible-changes.md` to describe the 
incompatibility caused by this PR.
   * [x] If you are contributing the connector code, please check that the 
following files are updated:
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to