wank125 opened a new issue, #18677:
URL: https://github.com/apache/dolphinscheduler/issues/18677

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/dolphinscheduler/issues?q=is%3Aissue) and 
found no similar issues.
   
   ### What happened
   
   When a workflow definition is created via the **export → import** round trip 
(`POST /projects/{projectCode}/workflow-definition/import`), the imported task 
definitions can have `taskExecuteType = null` (the exported JSON carries no 
`taskExecuteType` and the import path does not default it). The column 
`t_ds_task_definition.task_execute_type` has no default value, so `null` is 
persisted.
   
   When such a task finishes executing, the **task finish event handling chain 
crashes with an NPE and the task instance is stuck in RUNNING forever**:
   
   ```
   at 
org.apache.dolphinscheduler.server.master.utils.WorkflowInstanceUtils.logTaskInstanceInDetail(WorkflowInstanceUtils.java:94)
   at 
org.apache.dolphinscheduler.server.master.runner.WorkflowExecuteRunnable.taskFinished(...)
   at 
org.apache.dolphinscheduler.server.master.event.TaskStateEventHandler.handleStateEvent(...)
   ```
   
   The line is a **pure logging helper**:
   
   ```java
   .append("Task Execute Type:      
").append(taskInstance.getTaskExecuteType().getDesc())
   ```
   
   Observed on 3.2.2 (the same line exists on current dev), the full impact 
chain (all from server logs):
   
   1. Worker finishes the shell and reports SUCCESS — the event **arrives** at 
the master (`Handle task instance state event, the current task instance state 
SUCCESS`)
   2. `taskFinished` → logging helper NPE
   3. The event is retried 4 times, all fail with the same NPE, then is dropped
   4. `t_ds_task_instance.state` stays RUNNING forever, the workflow instance 
never completes, downstream tasks never start
   5. The master keeps re-dispatching the shell (we observed the same task's 
shell log file created 4 times), i.e. **non-idempotent tasks would be executed 
repeatedly**
   
   In our environment 48 imported task definitions (46 tasks of one workflow + 
2 of another) all had `task_execute_type = NULL`, and every workflow whose task 
chain contains such a task reproduced the stuck state deterministically. 
Workflows created via the metadata-DB insert path (with the value set) never 
reproduce it. Fixing the data (`UPDATE t_ds_task_definition SET 
task_execute_type = 0 WHERE task_execute_type IS NULL`, same for 
`t_ds_task_definition_log`) immediately restores normal completion — which 
confirms the null `taskExecuteType` as the trigger.
   
   ### What you expected to happen
   
   1. A logging helper must not be able to crash the task-finish event chain — 
it should render a null `taskExecuteType` safely.
   2. The import path should default a missing `taskExecuteType` to `BATCH` 
instead of persisting `null` (the export → import round trip is the official 
path and should not produce data the runtime cannot handle).
   
   ### How to reproduce
   
   1. Export any workflow via `POST 
/projects/{projectCode}/workflow-definition/batch-export` (returns plain JSON)
   2. Import it back via `POST 
/projects/{projectCode}/workflow-definition/import` (multipart file, JSON 
array) — note the task definitions now have `task_execute_type = NULL` in 
`t_ds_task_definition`
   3. Run the workflow; when a task (typically one with an upstream dependency) 
finishes: NPE in the master log at 
`WorkflowInstanceUtils.logTaskInstanceInDetail`, task stuck in RUNNING, 
workflow instance never completes, shell re-dispatched repeatedly
   
   ### Anything else
   
   Fix proposal (happy to submit the PR): null-guard in 
`logTaskInstanceInDetail` + default `taskExecuteType = BATCH` in 
`ProcessServiceImpl#saveTaskDefine` so both the persisted data and the runtime 
are safe. Possibly also worth adding a default value to the `task_execute_type` 
column for new schemas.
   
   ### Version
   
   3.2.2 (line 74 there), and the same unguarded line exists on current dev 
(line 94).
   
   ### Are you willing to submit PR?
   
   - [x] Yes I am willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: 
[email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to