DanielLeens opened a new issue, #11667:
URL: https://github.com/apache/seatunnel/issues/11667

   ## Background
   SeaTunnel already displays the current error on the job page, but only as a 
single raw string: `job.errorMsg` is rendered as one unstructured `<pre>` block 
(`seatunnel-engine-ui/src/views/jobs/detail.tsx`), with no per-task attribution 
and no history. Internally the story is the same: `PhysicalPlan.errorBySubPlan` 
is an `AtomicReference<String>` that uses `compareAndSet(null, ...)` to capture 
only the *first* error for a sub-plan and silently discards every error after 
that (`seatunnel-engine/.../dag/physical/PhysicalPlan.java`). 
`TaskExecutionState` does carry a `throwableMsg` per status report, but it is 
transient and never persisted per attempt, and there is currently no 
restart/attempt counter anywhere in the engine 
(`restartCount`/`attemptNumber`/`retryCount` do not exist).
   
   Closely related: every pipeline and task group already maintains a 
timestamped state-transition history (CREATED to DEPLOYING to RUNNING to 
FINISHED/CANCELING/CANCELED/FAILING/FAILED) in a `stateTimestamps` array backed 
by a shared internal map (`SubPlan.java`, `PhysicalVertex.java`), but only the 
**job-level** timestamps are ever read back (`JobMaster#getStateTimestamp`); 
the pipeline- and task-group-level timestamps are recorded and then never 
exposed anywhere. That is exactly the data a "when did this task start failing, 
and how does that compare to its previous attempt" view would need, and it 
already exists in memory today.
   
   Flink's Web UI addresses both needs with its "Exceptions" tab: the root 
exception plus the full exception history across restarts, each entry 
timestamped and attributed to a task/host.
   
   ## Problem to solve
   Because only the first error per sub-plan is kept and there is no attempt 
concept, diagnosing an intermittently-failing job means either catching the 
error in real time before something else overwrites the single stored value, or 
digging through raw logs across every affected worker by hand. This makes it 
hard to tell a one-off failure apart from a recurring pattern.
   
   ## Proposed scope
   Add a dedicated exception/failure history view, scoped to a job:
   - persist more than the first error per sub-plan/task: keep a bounded 
history of failures (timestamp, failing task/subtask, host/worker, exception 
type and message)
   - introduce an attempt/restart counter so history entries can be grouped by 
attempt, since none exists today
   - expose the already-tracked-but-unread pipeline/task-group 
`stateTimestamps` alongside each failure entry, so a failure can be shown 
against how long that attempt had been running rather than as an isolated event
   - apply to both running and finished jobs, within Zeta's existing history 
retention window
   - link each entry to the corresponding task's log (building on #9050/#11662)
   
   ## Why this should be a dedicated feature
   This requires retaining more than "the current error," which has retention, 
storage, and API-contract implications:
   - how many historical failures are kept per job, and for how long after the 
job finishes
   - how this interacts with existing finished-job history storage (including 
the known S3-backed finished-jobs issue, #10039) so it does not regress at scale
   - exposing the pipeline/task-group `stateTimestamps` is extending an 
existing internal model to a new read path, not new data collection, but it 
still needs a defined contract
   
   ## STIP requirement before implementation
   Because this changes what failure data Zeta retains and exposes as a stable 
contract, **the claimant should submit a STIP design first and get maintainer 
agreement before starting implementation**. The STIP should clarify:
   - the exception-history data model, its retention limits for running vs. 
finished jobs, and how attempts are numbered
   - how large exception messages/stack traces are truncated or paginated
   - how this scales for jobs stored in external history backends (e.g. S3), 
consistent with the constraints already surfaced in #10039
   - which fields are stable API/UI contract versus best-effort telemetry
   
   ## Acceptance criteria
   - Users can see the full sequence of failures/restarts for a job, not only 
the latest one, from both the running and finished job views.
   - Each entry identifies the failing task/attempt/host and links to its log.
   - Recurring root causes across attempts are visually distinguishable from 
one-off failures.
   - English and Chinese docs are updated.
   
   ## Non-goals for the first version
   - automatic root-cause classification or suggested fixes
   - indefinite retention of full stack traces for all historical jobs
   - cross-job failure correlation/analytics
   
   ## Related work
   - Issue #9105 (closed) - exception message formatting
   - Issue #9050 (closed) - worker log file viewing
   - Issue #11662 - job log links currently broken
   - Issue #10039 (closed) - finished-jobs history at scale on S3
   - Issue #11351
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to