The GitHub Actions job "Required Checks" on texera.git/main has succeeded.
Run started by GitHub user github-merge-queue[bot] (triggered by 
github-merge-queue[bot]).

Head commit for run:
3733b2a11948a9ed83db58864db980dc4d7d6bc6 / Ryan Zhang 
<[email protected]>
feat(notebook-migration): support migrating Python files into workflows (#8544)

### What changes were proposed in this PR?

The migration tool accepted only a Jupyter notebook. It now accepts a
Python file as well, and produces the same result: a workflow, a stored
source file, and an operator mapping that drives highlighting in the
Jupyter panel.

The obstacle is that a notebook arrives already split into cells, and
those cells are the join key for the mapping. A script has no such
boundaries. Rather than guess at them with a heuristic splitter, this PR
asks the model for them: the script is sent with a line number on each
line, the model reports which line ranges became which UDF, and the
cells are derived from that answer. The model is already deciding which
code becomes which operator, so it is the right thing to ask.

Everything after the model's reply is shared with the notebook path.
Same mapping shape, same storage, same Jupyter upload, same highlight
index. No backend, database, Helm, or dependency changes.

Main pieces:

* `script-segmentation.ts`, a pure module that reconciles the reported
ranges and derives cells. Ranges can arrive reversed, overlapping, out
of order, past the end of the file, or not at all. Every case is
reconciled rather than rejected, because the ranges come back alongside
a workflow that already cost a full conversion. It cuts the file at the
union of all reported boundaries, which makes overlaps and gaps fall out
of one mechanism: an overlap becomes a shared cell mapped to both UDFs,
and a gap becomes a cell no UDF claims, so no source is ever lost.
* Script variants of the conversion prompt and of the worked example in
the documentation prelude. The existing worked example is written in `#
START CELL1` form and seeded as a system message, so a script conversion
needed its own, or the model would answer in cell ids for an input that
has none.

`migration-prompts.ts` is 175 insertions and 1 deletion. All but five of
those lines are the new script constants. The notebook prompts take two
changes, both fixes to pre-existing bugs found during review:
`MAPPING_PROMPT`'s example mapped only four of the five UDFs its worked
example defines, which taught the model that leaving a UDF unmapped is
acceptable, and `WORKFLOW_PROMPT` carried a malformed sentence ("every
distinct UDF that uses that constructs an object of that class") that
left the class duplication rule ambiguous. Nothing else in the notebook
path changed.
* `convertScriptToWorkflow` on the LLM client, plus `parseScriptFile`
and `sendScriptToAIGenerateWorkflow` on the service.
* A second tab in the AI generate modal, and extension dispatch in the
dashboard entry point.

Two behavior-neutral refactors support this: workflow assembly and
mapping inversion were extracted out of `convertNotebookToWorkflow`, and
the LLM session lifecycle (initialize, verify, close in a `finally`) was
extracted so both paths share it and neither can leak a session.

One deliberate behavior change rides along. The shared mapping inversion
now skips a UDF whose cell list is not an array, with a warning, rather
than throwing. A model that answers with a bare string there previously
discarded an entire conversion after two expensive calls; it now
degrades the mapping instead, which is what the script path already does
for a malformed range.

#### Demo video
(waiting was cut out to keep demo short)


https://github.com/user-attachments/assets/2e248258-4520-4117-b67b-d1e10ed5bd5c

### Any related issues, documentation, discussions?

Closes #8007
Parent-issue #4301.

The choice to keep the mapping and Jupyter panel for scripts, rather
than returning only a workflow, was discussed in
https://github.com/apache/texera/discussions/8489.

The panel and toolbar copy still says "Jupyter Notebook" where a script
user would expect something general; that is being handled in #8543 so
this PR stays behavioral.

### How was this PR tested?

Unit tests, all new unless noted:

* `script-segmentation.spec.ts`: 32 tests over well-formed input, gaps,
overlaps, malformed and out of bounds ranges, the accepted range shapes,
degenerate input, and source fidelity. One test asserts that every line
with content lands in exactly one cell, in order, which is the property
that protects users from losing code. A blank run isolated between two
reported ranges is dropped on purpose, so the cells do not reassemble
byte for byte.
* `migration-llm.spec.ts`: 31 to 42. Covers line numbering, prelude
selection, cell derivation, the notebook it returns, degrading to one
unmapped cell when no usable ranges come back, and the non-array mapping
entry described above.
* `notebook-migration.service.spec.ts`: 39 to 48. Lifecycle and
`parseScriptFile`.
* `notebook-import-modal.component.spec.ts`: 19 to 27. Tab switching,
per tab upload targets, and submission.
* `user-workflow.component.spec.ts`: 76 to 81. Extension dispatch, and
that a `.py` stores the derived notebook and never reaches the notebook
path.

Full frontend suite passes (218 files, 5991 tests), along with `yarn
format:ci` and a production build.

Manually tested end to end against a live model: uploaded a `.py`,
confirmed the generated workflow, confirmed the derived notebook opens
in the Jupyter panel, and confirmed clicking an operator highlights the
cell its code came from.

### Was this PR authored or co-authored using generative AI tooling?

Generated-by: Claude Code (Claude Opus 5)

Report URL: https://github.com/apache/texera/actions/runs/35154569591

With regards,
GitHub Actions via GitBox

Reply via email to