This is an automated email from the ASF dual-hosted git repository.
mattcasters pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/hop.git
The following commit(s) were added to refs/heads/main by this push:
new a7ff631d7a Fixes #8640 : Explain pipelines and workflows (#8645)
a7ff631d7a is described below
commit a7ff631d7a150040929a1a641438d15e539eadd8
Author: Matt Casters <[email protected]>
AuthorDate: Mon Sep 28 12:25:26 2026 +0200
Fixes #8640 : Explain pipelines and workflows (#8645)
* Fixes #8640 : Explain pipelines and workflows with interactive
infographics and detailed conceptual guides
* Fixes #8640 : address PR review feedback
- Add the missing blank line before eight lists that Asciidoctor was
folding back into the preceding paragraph
- Follow the site's [data-theme] toggle in the three infographics, not
just the OS prefers-color-scheme setting
- Correct the load-balancing engine descriptions: each run is assigned to
one server from a group, with retry when that server is at capacity
- Fix the overlapping title and the clipped caption in the workflow diagram
- Replace hardcoded text fills that dark mode could not override
- Scope the back-pressure memory claim to the buffers between transforms
- Prefix the SVG CSS classes and keyframes per diagram so the three
inlined style blocks stop colliding
---------
Co-authored-by: Bart Maertens <[email protected]>
---
.../images/concepts/hop-architecture-overview.svg | 186 +++++++++++++++++
.../images/concepts/pipeline-parallel-flow.svg | 229 +++++++++++++++++++++
.../images/concepts/workflow-backtracking-flow.svg | 226 ++++++++++++++++++++
.../modules/ROOT/pages/concepts.adoc | 15 +-
.../ROOT/pages/getting-started/hop-concepts.adoc | 63 ++++--
.../pages/getting-started/hop-gui-pipelines.adoc | 21 +-
.../pages/getting-started/hop-gui-workflows.adoc | 24 ++-
.../modules/ROOT/pages/pipeline/pipelines.adoc | 177 +++++++++++++---
.../ROOT/pages/snippets/hop-concepts/action.adoc | 9 +-
.../pages/snippets/hop-concepts/item-types.adoc | 6 +-
.../ROOT/pages/snippets/hop-concepts/pipeline.adoc | 9 +-
.../pages/snippets/hop-concepts/transform.adoc | 10 +-
.../ROOT/pages/snippets/hop-concepts/workflow.adoc | 9 +-
.../modules/ROOT/pages/workflow/workflows.adoc | 154 +++++++++++---
14 files changed, 1028 insertions(+), 110 deletions(-)
diff --git
a/docs/hop-user-manual/modules/ROOT/assets/images/concepts/hop-architecture-overview.svg
b/docs/hop-user-manual/modules/ROOT/assets/images/concepts/hop-architecture-overview.svg
new file mode 100644
index 0000000000..8ad97c842a
--- /dev/null
+++
b/docs/hop-user-manual/modules/ROOT/assets/images/concepts/hop-architecture-overview.svg
@@ -0,0 +1,186 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<!--
+ Licensed to the Apache Software Foundation (ASF) under one
+ or more contributor license agreements. See the NOTICE file
+ distributed with this work for additional information
+ regarding copyright ownership. The ASF licenses this file
+ to you under the Apache License, Version 2.0 (the
+ "License"); you may not use this file except in compliance
+ with the License. You may obtain a copy of the License at
+
+ http://www.apache.org/licenses/LICENSE-2.0
+
+ Unless required by applicable law or agreed to in writing,
+ software distributed under the License is distributed on an
+ "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+ KIND, either express or implied. See the License for the
+ specific language governing permissions and limitations
+ under the License.
+-->
+<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 480" width="960"
height="480">
+ <defs>
+ <style>
+ .ar-bg { fill: #f8fafc; }
+ .ar-layer-box { fill: #ffffff; stroke: #cbd5e1; stroke-width: 1.5; rx:
8px; }
+ .ar-layer-header-wf { fill: #0e3a5a; }
+ .ar-layer-header-pl { fill: #00a3a6; }
+ .ar-layer-header-conn { fill: #334155; }
+
+ .ar-text-head { font-family: system-ui, -apple-system, sans-serif;
font-size: 14px; font-weight: 700; fill: #ffffff; }
+ .ar-text-item-title { font-family: system-ui, -apple-system, sans-serif;
font-size: 13px; font-weight: 700; fill: #0e3a5a; }
+ .ar-text-item-desc { font-family: system-ui, -apple-system, sans-serif;
font-size: 11px; fill: #64748b; }
+
+ .ar-pill { fill: #f1f5f9; stroke: #cbd5e1; stroke-width: 1; rx: 6px; }
+ .ar-pill-accent { fill: #e0f2fe; stroke: #bae6fd; stroke-width: 1; rx:
6px; }
+ .ar-pill-text { font-family: system-ui, -apple-system, sans-serif;
font-size: 11px; font-weight: 600; fill: #0369a1; }
+
+ .ar-flow-arrow { stroke: #00a3a6; stroke-width: 2; stroke-dasharray: 4
4; fill: none; }
+ .ar-ink-strong { fill: #0e3a5a; }
+
+ /* Dark mode. The OS preference is the default, but the Hop site also has
+ its own light/dark toggle which sets data-theme on the page root, so
+ both have to be honoured -- and an explicit light choice has to win
+ over a dark OS. */
+ @media (prefers-color-scheme: dark) {
+ :root:not([data-theme="light"]) .ar-bg { fill: #0f172a; }
+ :root:not([data-theme="light"]) .ar-layer-box { fill: #1e293b; stroke:
#334155; }
+ :root:not([data-theme="light"]) .ar-text-item-title { fill: #f1f5f9; }
+ :root:not([data-theme="light"]) .ar-text-item-desc { fill: #94a3b8; }
+ :root:not([data-theme="light"]) .ar-pill { fill: #0f172a; stroke:
#334155; }
+ :root:not([data-theme="light"]) .ar-pill-accent { fill: #0e3a5a;
stroke: #0284c7; }
+ :root:not([data-theme="light"]) .ar-pill-text { fill: #38bdf8; }
+ :root:not([data-theme="light"]) .ar-ink-strong { fill: #f1f5f9; }
+ }
+
+ :root[data-theme="dark"] .ar-bg { fill: #0f172a; }
+ :root[data-theme="dark"] .ar-layer-box { fill: #1e293b; stroke: #334155;
}
+ :root[data-theme="dark"] .ar-text-item-title { fill: #f1f5f9; }
+ :root[data-theme="dark"] .ar-text-item-desc { fill: #94a3b8; }
+ :root[data-theme="dark"] .ar-pill { fill: #0f172a; stroke: #334155; }
+ :root[data-theme="dark"] .ar-pill-accent { fill: #0e3a5a; stroke:
#0284c7; }
+ :root[data-theme="dark"] .ar-pill-text { fill: #38bdf8; }
+ :root[data-theme="dark"] .ar-ink-strong { fill: #f1f5f9; }
+ </style>
+
+ <marker id="arrow-down" viewBox="0 0 10 10" refX="5" refY="6"
markerWidth="6" markerHeight="6" orient="auto">
+ <path d="M 0 0 L 10 0 L 5 8 z" fill="#00a3a6" />
+ </marker>
+ </defs>
+
+ <!-- Background -->
+ <rect width="960" height="480" rx="10" class="ar-bg" />
+
+ <!-- TOP BAR: CLIENT INTERFACES & SCHEDULERS -->
+ <g transform="translate(30, 15)">
+ <rect width="900" height="52" class="ar-layer-box" />
+ <text x="20" y="31" class="ar-text-item-title" font-size="13">Clients
& Schedulers:</text>
+
+ <g transform="translate(180, 10)">
+ <rect width="135" height="32" class="ar-pill" />
+ <text x="67" y="21" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="11.5" font-weight="600"
class="ar-ink-strong">Hop GUI</text>
+ </g>
+ <g transform="translate(325, 10)">
+ <rect width="135" height="32" class="ar-pill" />
+ <text x="67" y="21" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="11.5" font-weight="600"
class="ar-ink-strong">hop-run (CLI)</text>
+ </g>
+ <g transform="translate(470, 10)">
+ <rect width="145" height="32" class="ar-pill" />
+ <text x="72" y="21" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="11.5" font-weight="600"
class="ar-ink-strong">Python (PyHop)</text>
+ </g>
+ <g transform="translate(625, 10)">
+ <rect width="250" height="32" class="ar-pill-accent" />
+ <text x="125" y="21" text-anchor="middle" class="ar-pill-text">Airflow /
Cron / Kubernetes</text>
+ </g>
+ </g>
+
+ <!-- DOWN ARROW 1 -->
+ <line x1="480" y1="68" x2="480" y2="84" class="ar-flow-arrow"
marker-end="url(#arrow-down)" />
+
+ <!-- LAYER 1: WORKFLOWS (ORCHESTRATOR / CONTROL PLANE) -->
+ <g transform="translate(30, 85)">
+ <rect width="900" height="120" class="ar-layer-box" />
+ <path d="M 0 8 C 0 3.58 3.58 0 8 0 L 892 0 C 896.42 0 900 3.58 900 8 L 900
32 L 0 32 Z" class="ar-layer-header-wf" />
+ <text x="16" y="21" class="ar-text-head">Workflows (.hwf) — The Control
Plane (Sequential Orchestration)</text>
+
+ <!-- Actions Inside Workflow -->
+ <g transform="translate(16, 42)">
+ <rect width="265" height="66" class="ar-pill" />
+ <text x="14" y="24" class="ar-text-item-title">Prerequisites &
Health</text>
+ <text x="14" y="42" class="ar-text-item-desc">Database checks, file
verification,</text>
+ <text x="14" y="56" class="ar-text-item-desc">archive extraction, PGP
encryption</text>
+ </g>
+
+ <g transform="translate(297, 42)">
+ <rect width="285" height="66" class="ar-pill" />
+ <text x="14" y="24" class="ar-text-item-title">Decision Routing &
Backtracking</text>
+ <text x="14" y="42" class="ar-text-item-desc">Success hops, failure
hops, error alerts,</text>
+ <text x="14" y="56" class="ar-text-item-desc">conditional recovery
paths, parallel branches</text>
+ </g>
+
+ <g transform="translate(598, 42)">
+ <rect width="286" height="66" class="ar-pill-accent" />
+ <text x="14" y="24" class="ar-pill-text" font-size="12">Workflow
Engines</text>
+ <text x="14" y="42" class="ar-text-item-desc">Native Local (Workstation
/ Server)</text>
+ <text x="14" y="56" class="ar-text-item-desc">Native Remote (Hop Server)
& Clustered</text>
+ </g>
+ </g>
+
+ <!-- DOWN ARROW 2 -->
+ <line x1="480" y1="206" x2="480" y2="222" class="ar-flow-arrow"
marker-end="url(#arrow-down)" />
+
+ <!-- LAYER 2: PIPELINES (DATA WORKER / DATA PLANE) -->
+ <g transform="translate(30, 225)">
+ <rect width="900" height="120" class="ar-layer-box" />
+ <path d="M 0 8 C 0 3.58 3.58 0 8 0 L 892 0 C 896.42 0 900 3.58 900 8 L 900
32 L 0 32 Z" class="ar-layer-header-pl" />
+ <text x="16" y="21" class="ar-text-head">Pipelines (.hpl) — The Data Plane
(Parallel Streaming Data Workers)</text>
+
+ <!-- Content Inside Pipeline Layer -->
+ <g transform="translate(16, 42)">
+ <rect width="265" height="66" class="ar-pill" />
+ <text x="14" y="24" class="ar-text-item-title">Streaming
Transforms</text>
+ <text x="14" y="42" class="ar-text-item-desc">Reads from sources,
cleans, transforms,</text>
+ <text x="14" y="56" class="ar-text-item-desc">enriches, joins, and
writes simultaneously</text>
+ </g>
+
+ <g transform="translate(297, 42)">
+ <rect width="285" height="66" class="ar-pill" />
+ <text x="14" y="24" class="ar-text-item-title">Row Buffers &
Back-Pressure</text>
+ <text x="14" y="42" class="ar-text-item-desc">In-memory queues stream
rows safely</text>
+ <text x="14" y="56" class="ar-text-item-desc">Slow downstream targets
throttle readers</text>
+ </g>
+
+ <g transform="translate(598, 42)">
+ <rect width="286" height="66" class="ar-pill-accent" />
+ <text x="14" y="24" class="ar-pill-text" font-size="12">Pipeline
Execution Engines</text>
+ <text x="14" y="42" class="ar-text-item-desc">Native Local, Remote Hop
Server, Beam</text>
+ <text x="14" y="56" class="ar-text-item-desc">(Flink, Spark, Dataflow),
Native Spark</text>
+ </g>
+ </g>
+
+ <!-- DOWN ARROW 3 -->
+ <line x1="480" y1="346" x2="480" y2="362" class="ar-flow-arrow"
marker-end="url(#arrow-down)" />
+
+ <!-- LAYER 3: STORAGE, PLATFORMS & CONNECTIVITY -->
+ <g transform="translate(30, 365)">
+ <rect width="900" height="95" class="ar-layer-box" />
+ <path d="M 0 8 C 0 3.58 3.58 0 8 0 L 892 0 C 896.42 0 900 3.58 900 8 L 900
28 L 0 28 Z" class="ar-layer-header-conn" />
+ <text x="16" y="19" class="ar-text-head" font-size="12.5">Storage &
External Platforms (via VFS & Connectors)</text>
+
+ <g transform="translate(16, 38)">
+ <rect width="200" height="44" class="ar-pill" />
+ <text x="100" y="27" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="12" font-weight="600"
class="ar-ink-strong">Relational Databases</text>
+ </g>
+ <g transform="translate(232, 38)">
+ <rect width="215" height="44" class="ar-pill" />
+ <text x="107" y="27" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="12" font-weight="600"
class="ar-ink-strong">Cloud Storage (S3/GCS/Blob)</text>
+ </g>
+ <g transform="translate(463, 38)">
+ <rect width="215" height="44" class="ar-pill" />
+ <text x="107" y="27" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="12" font-weight="600"
class="ar-ink-strong">Streaming (Kafka/PubSub)</text>
+ </g>
+ <g transform="translate(694, 38)">
+ <rect width="190" height="44" class="ar-pill" />
+ <text x="95" y="27" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="12" font-weight="600"
class="ar-ink-strong">Data Warehouses & Lakes</text>
+ </g>
+ </g>
+</svg>
diff --git
a/docs/hop-user-manual/modules/ROOT/assets/images/concepts/pipeline-parallel-flow.svg
b/docs/hop-user-manual/modules/ROOT/assets/images/concepts/pipeline-parallel-flow.svg
new file mode 100644
index 0000000000..3b5e330924
--- /dev/null
+++
b/docs/hop-user-manual/modules/ROOT/assets/images/concepts/pipeline-parallel-flow.svg
@@ -0,0 +1,229 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<!--
+ Licensed to the Apache Software Foundation (ASF) under one
+ or more contributor license agreements. See the NOTICE file
+ distributed with this work for additional information
+ regarding copyright ownership. The ASF licenses this file
+ to you under the Apache License, Version 2.0 (the
+ "License"); you may not use this file except in compliance
+ with the License. You may obtain a copy of the License at
+
+ http://www.apache.org/licenses/LICENSE-2.0
+
+ Unless required by applicable law or agreed to in writing,
+ software distributed under the License is distributed on an
+ "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+ KIND, either express or implied. See the License for the
+ specific language governing permissions and limitations
+ under the License.
+-->
+<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 280" width="880"
height="280">
+ <defs>
+ <style>
+ .pl-bg { fill: #f8fafc; }
+ .pl-card { fill: #ffffff; stroke: #cbd5e1; stroke-width: 1.5; rx: 8px; }
+ .pl-card-source { fill: #ffffff; stroke: #00a3a6; stroke-width: 2.5; rx:
8px; }
+ .pl-card-target { fill: #ffffff; stroke: #f59e0b; stroke-width: 2.5; rx:
8px; }
+
+ .pl-text-title { font-family: system-ui, -apple-system, sans-serif;
font-size: 15px; font-weight: 700; fill: #0e3a5a; }
+ .pl-text-subtitle { font-family: system-ui, -apple-system, sans-serif;
font-size: 12px; fill: #64748b; }
+ .pl-text-tag { font-family: system-ui, -apple-system, sans-serif;
font-size: 10.5px; font-weight: 700; text-transform: uppercase; }
+ .pl-text-metric { font-family: system-ui, -apple-system, sans-serif;
font-size: 12px; font-weight: 700; }
+ .pl-text-buffer-title { font-family: system-ui, -apple-system,
sans-serif; font-size: 12.5px; font-weight: 700; fill: #0e3a5a; }
+
+ .pl-hop-line { stroke: #94a3b8; stroke-width: 3; stroke-dasharray: 6 4;
fill: none; }
+ .pl-buffer-box { fill: #f8fafc; stroke: #cbd5e1; stroke-width: 1.5; rx:
8px; }
+ .pl-bar-normal { fill: #00a3a6; rx: 2px; }
+ .pl-bar-warn { fill: #f59e0b; rx: 2px; }
+ .pl-bar-empty { fill: #e2e8f0; rx: 2px; }
+ .pl-ribbon { fill: #ffffff; stroke: #cbd5e1; stroke-width: 1; rx: 6px; }
+ .pl-ribbon-text { font-family: system-ui, -apple-system, sans-serif;
font-size: 12px; font-weight: 600; fill: #334155; }
+
+ .pl-pulse-warn { animation: pl-pulseWarning 2s infinite ease-in-out; }
+
+ @keyframes pl-pulseWarning {
+ 0%, 100% { filter: drop-shadow(0 0 2px rgba(245, 158, 11, 0.4)); }
+ 50% { filter: drop-shadow(0 0 8px rgba(245, 158, 11, 0.8)); }
+ }
+
+ /* Animated streaming row tokens */
+ .pl-dot-stream1 { animation: pl-streamAnim1 2s infinite linear; }
+ .pl-dot-stream2 { animation: pl-streamAnim2 2s infinite linear; }
+
+ @keyframes pl-streamAnim1 {
+ 0% { transform: translateX(0px); opacity: 0; }
+ 15% { opacity: 1; }
+ 85% { opacity: 1; }
+ 100% { transform: translateX(160px); opacity: 0; }
+ }
+
+ @keyframes pl-streamAnim2 {
+ 0% { transform: translateX(0px); opacity: 0; }
+ 15% { opacity: 1; }
+ 85% { opacity: 1; }
+ 100% { transform: translateX(160px); opacity: 0; }
+ }
+ .pl-ink-muted { fill: #475569; }
+
+ /* Dark mode. The OS preference is the default, but the Hop site also has
+ its own light/dark toggle which sets data-theme on the page root, so
+ both have to be honoured -- and an explicit light choice has to win
+ over a dark OS. */
+ @media (prefers-color-scheme: dark) {
+ :root:not([data-theme="light"]) .pl-bg { fill: #0f172a; }
+ :root:not([data-theme="light"]) .pl-card { fill: #1e293b; stroke:
#334155; }
+ :root:not([data-theme="light"]) .pl-card-source { fill: #1e293b;
stroke: #00a3a6; }
+ :root:not([data-theme="light"]) .pl-card-target { fill: #1e293b;
stroke: #f59e0b; }
+ :root:not([data-theme="light"]) .pl-text-title { fill: #f1f5f9; }
+ :root:not([data-theme="light"]) .pl-text-subtitle { fill: #94a3b8; }
+ :root:not([data-theme="light"]) .pl-text-buffer-title { fill: #cbd5e1;
}
+ :root:not([data-theme="light"]) .pl-buffer-box { fill: #0f172a;
stroke: #334155; }
+ :root:not([data-theme="light"]) .pl-bar-empty { fill: #334155; }
+ :root:not([data-theme="light"]) .pl-hop-line { stroke: #475569; }
+ :root:not([data-theme="light"]) .pl-ribbon { fill: #1e293b; stroke:
#334155; }
+ :root:not([data-theme="light"]) .pl-ribbon-text { fill: #cbd5e1; }
+ :root:not([data-theme="light"]) .pl-ink-muted { fill: #94a3b8; }
+ }
+
+ :root[data-theme="dark"] .pl-bg { fill: #0f172a; }
+ :root[data-theme="dark"] .pl-card { fill: #1e293b; stroke: #334155; }
+ :root[data-theme="dark"] .pl-card-source { fill: #1e293b; stroke:
#00a3a6; }
+ :root[data-theme="dark"] .pl-card-target { fill: #1e293b; stroke:
#f59e0b; }
+ :root[data-theme="dark"] .pl-text-title { fill: #f1f5f9; }
+ :root[data-theme="dark"] .pl-text-subtitle { fill: #94a3b8; }
+ :root[data-theme="dark"] .pl-text-buffer-title { fill: #cbd5e1; }
+ :root[data-theme="dark"] .pl-buffer-box { fill: #0f172a; stroke:
#334155; }
+ :root[data-theme="dark"] .pl-bar-empty { fill: #334155; }
+ :root[data-theme="dark"] .pl-hop-line { stroke: #475569; }
+ :root[data-theme="dark"] .pl-ribbon { fill: #1e293b; stroke: #334155; }
+ :root[data-theme="dark"] .pl-ribbon-text { fill: #cbd5e1; }
+ :root[data-theme="dark"] .pl-ink-muted { fill: #94a3b8; }
+ </style>
+
+ <marker id="arrow" viewBox="0 0 10 10" refX="6" refY="5" markerWidth="6"
markerHeight="6" orient="auto-start-reverse">
+ <path d="M 0 1 L 8 5 L 0 9 z" fill="#00a3a6" />
+ </marker>
+ </defs>
+
+ <!-- Background Canvas -->
+ <rect width="880" height="280" rx="10" class="pl-bg" />
+
+ <!-- Top Title Banner -->
+ <g transform="translate(25, 16)">
+ <rect width="830" height="36" rx="6" fill="#0e3a5a" />
+ <circle cx="20" cy="18" r="6" fill="#00a3a6" />
+ <text x="36" y="23" font-family="system-ui, -apple-system, sans-serif"
font-size="14.5" font-weight="700" fill="#ffffff">
+ Pipeline Streaming Execution Model
+ </text>
+ <text x="300" y="23" font-family="system-ui, -apple-system, sans-serif"
font-size="12.5" fill="#94a3b8">
+ — All transforms run concurrently; rows stream through in-memory buffers
+ </text>
+ </g>
+
+ <!-- TRANSFORM 1: SOURCE -->
+ <g transform="translate(25, 75)">
+ <rect width="170" height="135" class="pl-card-source" />
+ <rect x="12" y="12" width="64" height="20" rx="4" fill="#e0f2fe" />
+ <text x="44" y="26" text-anchor="middle" class="pl-text-tag"
fill="#0284c7">SOURCE</text>
+ <circle cx="152" cy="22" r="5" fill="#10b981" />
+
+ <text x="14" y="60" class="pl-text-title">Table Input</text>
+ <text x="14" y="80" class="pl-text-subtitle">Reads customer rows</text>
+ <line x1="14" y1="96" x2="156" y2="96" stroke="#e2e8f0" stroke-width="1" />
+ <text x="14" y="118" class="pl-text-metric" fill="#00a3a6">~8,000
rows/s</text>
+ </g>
+
+ <!-- HOP 1 (Connects Transform 1 to Transform 2) -->
+ <g transform="translate(195, 142)">
+ <line x1="0" y1="0" x2="160" y2="0" class="pl-hop-line"
marker-end="url(#arrow)" />
+
+ <!-- Animated row tokens -->
+ <g class="pl-dot-stream1">
+ <circle cx="12" cy="0" r="5" fill="#00a3a6" />
+ <circle cx="52" cy="0" r="5" fill="#00a3a6" />
+ <circle cx="92" cy="0" r="5" fill="#00a3a6" />
+ <circle cx="132" cy="0" r="5" fill="#00a3a6" />
+ </g>
+
+ <!-- Row Buffer Container 1 (Centered on hop) -->
+ <g transform="translate(25, -54)">
+ <rect width="110" height="108" class="pl-buffer-box" />
+ <text x="55" y="20" text-anchor="middle"
class="pl-text-buffer-title">Row Buffer</text>
+ <text x="55" y="35" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="10.5" class="pl-ink-muted">Capacity:
10k</text>
+
+ <!-- Visual buffer level bars -->
+ <g transform="translate(15, 45)">
+ <rect x="0" y="0" width="12" height="26" class="pl-bar-normal" />
+ <rect x="17" y="0" width="12" height="26" class="pl-bar-normal" />
+ <rect x="34" y="0" width="12" height="26" class="pl-bar-normal" />
+ <rect x="51" y="0" width="12" height="26" class="pl-bar-empty" />
+ <rect x="68" y="0" width="12" height="26" class="pl-bar-empty" />
+ </g>
+ <text x="55" y="94" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="11.5" font-weight="700"
fill="#00a3a6">Flowing</text>
+ </g>
+ </g>
+
+ <!-- TRANSFORM 2: INTERMEDIATE -->
+ <g transform="translate(355, 75)">
+ <rect width="170" height="135" class="pl-card" />
+ <rect x="12" y="12" width="86" height="20" rx="4" fill="#f3e8ff" />
+ <text x="55" y="26" text-anchor="middle" class="pl-text-tag"
fill="#7e22ce">TRANSFORM</text>
+ <circle cx="152" cy="22" r="5" fill="#10b981" />
+
+ <text x="14" y="60" class="pl-text-title">Lookup & Clean</text>
+ <text x="14" y="80" class="pl-text-subtitle">Enriches &
validates</text>
+ <line x1="14" y1="96" x2="156" y2="96" stroke="#e2e8f0" stroke-width="1" />
+ <text x="14" y="118" class="pl-text-metric" fill="#7e22ce">Paused
(Throttled)</text>
+ </g>
+
+ <!-- HOP 2 (Connects Transform 2 to Transform 3) -->
+ <g transform="translate(525, 142)">
+ <line x1="0" y1="0" x2="160" y2="0" class="pl-hop-line"
marker-end="url(#arrow)" />
+
+ <!-- Animated row tokens (amber/throttled) -->
+ <g class="pl-dot-stream2">
+ <circle cx="15" cy="0" r="5" fill="#f59e0b" />
+ <circle cx="65" cy="0" r="5" fill="#f59e0b" />
+ <circle cx="115" cy="0" r="5" fill="#f59e0b" />
+ </g>
+
+ <!-- Row Buffer Container 2 (Full - Back-pressure) -->
+ <g transform="translate(25, -54)" class="pl-pulse-warn">
+ <rect width="110" height="108" class="pl-buffer-box" stroke="#f59e0b"
stroke-width="2.5" />
+ <text x="55" y="20" text-anchor="middle" class="pl-text-buffer-title"
fill="#d97706">Buffer Full!</text>
+ <text x="55" y="35" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="10.5" font-weight="600"
fill="#d97706">10k / 10k</text>
+
+ <!-- Visual buffer level bars (Full) -->
+ <g transform="translate(15, 45)">
+ <rect x="0" y="0" width="12" height="26" class="pl-bar-warn" />
+ <rect x="17" y="0" width="12" height="26" class="pl-bar-warn" />
+ <rect x="34" y="0" width="12" height="26" class="pl-bar-warn" />
+ <rect x="51" y="0" width="12" height="26" class="pl-bar-warn" />
+ <rect x="68" y="0" width="12" height="26" class="pl-bar-warn" />
+ </g>
+ <text x="55" y="94" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="11" font-weight="700"
fill="#d97706">Back-Pressure</text>
+ </g>
+ </g>
+
+ <!-- TRANSFORM 3: TARGET SINK -->
+ <g transform="translate(685, 75)">
+ <rect width="170" height="135" class="pl-card-target" />
+ <rect x="12" y="12" width="86" height="20" rx="4" fill="#fef3c7" />
+ <text x="55" y="26" text-anchor="middle" class="pl-text-tag"
fill="#b45309">TARGET SINK</text>
+ <circle cx="152" cy="22" r="5" fill="#10b981" />
+
+ <text x="14" y="60" class="pl-text-title">Database Insert</text>
+ <text x="14" y="80" class="pl-text-subtitle">Writes to warehouse</text>
+ <line x1="14" y1="96" x2="156" y2="96" stroke="#e2e8f0" stroke-width="1" />
+ <text x="14" y="118" class="pl-text-metric" fill="#b45309">~2,000
rows/s</text>
+ </g>
+
+ <!-- Bottom Explanatory Ribbon -->
+ <g transform="translate(25, 226)">
+ <rect width="830" height="38" class="pl-ribbon" />
+ <circle cx="20" cy="19" r="4.5" fill="#f59e0b" />
+ <text x="34" y="23" class="pl-ribbon-text">
+ Back-Pressure: When Database Insert is slower than Table Input, the row
buffer fills up and automatically pauses Table Input.
+ </text>
+ </g>
+</svg>
diff --git
a/docs/hop-user-manual/modules/ROOT/assets/images/concepts/workflow-backtracking-flow.svg
b/docs/hop-user-manual/modules/ROOT/assets/images/concepts/workflow-backtracking-flow.svg
new file mode 100644
index 0000000000..4b1ef46ebc
--- /dev/null
+++
b/docs/hop-user-manual/modules/ROOT/assets/images/concepts/workflow-backtracking-flow.svg
@@ -0,0 +1,226 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<!--
+ Licensed to the Apache Software Foundation (ASF) under one
+ or more contributor license agreements. See the NOTICE file
+ distributed with this work for additional information
+ regarding copyright ownership. The ASF licenses this file
+ to you under the Apache License, Version 2.0 (the
+ "License"); you may not use this file except in compliance
+ with the License. You may obtain a copy of the License at
+
+ http://www.apache.org/licenses/LICENSE-2.0
+
+ Unless required by applicable law or agreed to in writing,
+ software distributed under the License is distributed on an
+ "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+ KIND, either express or implied. See the License for the
+ specific language governing permissions and limitations
+ under the License.
+-->
+<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 310" width="880"
height="310">
+ <defs>
+ <style>
+ .wf-bg { fill: #f8fafc; }
+ .wf-card { fill: #ffffff; stroke: #cbd5e1; stroke-width: 1.5; rx: 8px; }
+ .wf-card-start { fill: #ffffff; stroke: #0e3a5a; stroke-width: 2.5; rx:
8px; }
+ .wf-card-success { fill: #ffffff; stroke: #10b981; stroke-width: 2.5;
rx: 8px; }
+ .wf-card-fail { fill: #ffffff; stroke: #ef4444; stroke-width: 2.5; rx:
8px; }
+
+ .wf-text-title { font-family: system-ui, -apple-system, sans-serif;
font-size: 14.5px; font-weight: 700; fill: #0e3a5a; }
+ .wf-text-subtitle { font-family: system-ui, -apple-system, sans-serif;
font-size: 12px; fill: #64748b; }
+ .wf-text-tag { font-family: system-ui, -apple-system, sans-serif;
font-size: 10.5px; font-weight: 700; text-transform: uppercase; }
+
+ .wf-hop-uncond { stroke: #0e3a5a; stroke-width: 2.5; fill: none; }
+ .wf-hop-success { stroke: #10b981; stroke-width: 2.5; fill: none; }
+ .wf-hop-failure { stroke: #ef4444; stroke-width: 2.5; stroke-dasharray:
6 3; fill: none; }
+
+ .wf-hop-label { font-family: system-ui, -apple-system, sans-serif;
font-size: 10.5px; font-weight: 700; }
+ .wf-ribbon { fill: #ffffff; stroke: #cbd5e1; stroke-width: 1; rx: 6px; }
+ .wf-ribbon-text { font-family: system-ui, -apple-system, sans-serif;
font-size: 12px; font-weight: 600; fill: #334155; }
+
+ /* Step token animation along success branch */
+ .wf-token-seq {
+ animation: wf-seqFollow 6s infinite ease-in-out;
+ }
+
+ @keyframes wf-seqFollow {
+ 0% { transform: translate(65px, 150px); opacity: 0; }
+ 5% { opacity: 1; }
+ 25% { transform: translate(220px, 150px); }
+ 35% { transform: translate(220px, 150px); }
+ 55% { transform: translate(450px, 105px); }
+ 65% { transform: translate(450px, 105px); }
+ 85% { transform: translate(670px, 105px); opacity: 1; }
+ 95% { transform: translate(805px, 105px); opacity: 1; }
+ 100% { transform: translate(805px, 105px); opacity: 0; }
+ }
+ .wf-ink-muted { fill: #475569; }
+ .wf-ink-strong { fill: #0e3a5a; }
+
+ /* Dark mode. The OS preference is the default, but the Hop site also has
+ its own light/dark toggle which sets data-theme on the page root, so
+ both have to be honoured -- and an explicit light choice has to win
+ over a dark OS. */
+ @media (prefers-color-scheme: dark) {
+ :root:not([data-theme="light"]) .wf-bg { fill: #0f172a; }
+ :root:not([data-theme="light"]) .wf-card { fill: #1e293b; stroke:
#334155; }
+ :root:not([data-theme="light"]) .wf-card-start { fill: #1e293b;
stroke: #60a5fa; }
+ :root:not([data-theme="light"]) .wf-card-success { fill: #1e293b;
stroke: #10b981; }
+ :root:not([data-theme="light"]) .wf-card-fail { fill: #1e293b; stroke:
#ef4444; }
+ :root:not([data-theme="light"]) .wf-text-title { fill: #f1f5f9; }
+ :root:not([data-theme="light"]) .wf-text-subtitle { fill: #94a3b8; }
+ :root:not([data-theme="light"]) .wf-hop-uncond { stroke: #94a3b8; }
+ :root:not([data-theme="light"]) .wf-ribbon { fill: #1e293b; stroke:
#334155; }
+ :root:not([data-theme="light"]) .wf-ribbon-text { fill: #cbd5e1; }
+ :root:not([data-theme="light"]) .wf-ink-muted { fill: #94a3b8; }
+ :root:not([data-theme="light"]) .wf-ink-strong { fill: #f1f5f9; }
+ }
+
+ :root[data-theme="dark"] .wf-bg { fill: #0f172a; }
+ :root[data-theme="dark"] .wf-card { fill: #1e293b; stroke: #334155; }
+ :root[data-theme="dark"] .wf-card-start { fill: #1e293b; stroke:
#60a5fa; }
+ :root[data-theme="dark"] .wf-card-success { fill: #1e293b; stroke:
#10b981; }
+ :root[data-theme="dark"] .wf-card-fail { fill: #1e293b; stroke: #ef4444;
}
+ :root[data-theme="dark"] .wf-text-title { fill: #f1f5f9; }
+ :root[data-theme="dark"] .wf-text-subtitle { fill: #94a3b8; }
+ :root[data-theme="dark"] .wf-hop-uncond { stroke: #94a3b8; }
+ :root[data-theme="dark"] .wf-ribbon { fill: #1e293b; stroke: #334155; }
+ :root[data-theme="dark"] .wf-ribbon-text { fill: #cbd5e1; }
+ :root[data-theme="dark"] .wf-ink-muted { fill: #94a3b8; }
+ :root[data-theme="dark"] .wf-ink-strong { fill: #f1f5f9; }
+ </style>
+
+ <marker id="arrow-uncond" viewBox="0 0 10 10" refX="6" refY="5"
markerWidth="6" markerHeight="6" orient="auto-start-reverse">
+ <path d="M 0 1 L 8 5 L 0 9 z" fill="#0e3a5a" />
+ </marker>
+ <marker id="arrow-succ" viewBox="0 0 10 10" refX="6" refY="5"
markerWidth="6" markerHeight="6" orient="auto-start-reverse">
+ <path d="M 0 1 L 8 5 L 0 9 z" fill="#10b981" />
+ </marker>
+ <marker id="arrow-fail" viewBox="0 0 10 10" refX="6" refY="5"
markerWidth="6" markerHeight="6" orient="auto-start-reverse">
+ <path d="M 0 1 L 8 5 L 0 9 z" fill="#ef4444" />
+ </marker>
+ </defs>
+
+ <!-- Background -->
+ <rect width="880" height="310" rx="10" class="wf-bg" />
+
+ <!-- Top Title Banner -->
+ <g transform="translate(25, 16)">
+ <rect width="830" height="36" rx="6" fill="#0e3a5a" />
+ <circle cx="20" cy="18" r="6" fill="#10b981" />
+ <text x="36" y="23" font-family="system-ui, -apple-system, sans-serif"
font-size="14.5" font-weight="700" fill="#ffffff">
+ Workflow Orchestration & Backtracking Model
+ </text>
+ <text x="392" y="23" font-family="system-ui, -apple-system, sans-serif"
font-size="12.5" fill="#94a3b8">
+ — One action at a time; hops route on success or failure
+ </text>
+ </g>
+
+ <!-- CONNECTIONS / HOPS -->
+
+ <!-- Hop 1: Start -> Check Files (Unconditional) -->
+ <line x1="105" y1="150" x2="165" y2="150" class="wf-hop-uncond"
marker-end="url(#arrow-uncond)" />
+ <rect x="110" y="131" width="52" height="18" rx="3" fill="#e2e8f0" />
+ <text x="136" y="144" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="9.5" font-weight="700"
class="wf-ink-strong">ALWAYS</text>
+
+ <!-- Hop 2 (Success Path): Check Files -> Run Pipeline -->
+ <path d="M 325 150 C 355 150, 360 105, 395 105" class="wf-hop-success"
marker-end="url(#arrow-succ)" />
+ <rect x="338" y="93" width="60" height="18" rx="3" fill="#d1fae5" />
+ <text x="368" y="106" text-anchor="middle" class="wf-hop-label"
fill="#059669">SUCCESS</text>
+
+ <!-- Hop 3: Run Pipeline -> Archive Files -->
+ <line x1="555" y1="105" x2="615" y2="105" class="wf-hop-success"
marker-end="url(#arrow-succ)" />
+ <rect x="565" y="90" width="42" height="17" rx="3" fill="#d1fae5" />
+ <text x="586" y="103" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="9.5" font-weight="700"
fill="#059669">TRUE</text>
+
+ <!-- Hop 4: Archive Files -> Success Action -->
+ <line x1="765" y1="105" x2="805" y2="105" class="wf-hop-success"
marker-end="url(#arrow-succ)" />
+
+ <!-- Hop 5 (Failure Path): Check Files -> Send Alert -->
+ <path d="M 325 150 C 355 150, 360 205, 395 205" class="wf-hop-failure"
marker-end="url(#arrow-fail)" />
+ <rect x="338" y="186" width="60" height="18" rx="3" fill="#fee2e2" />
+ <text x="368" y="199" text-anchor="middle" class="wf-hop-label"
fill="#dc2626">FAILURE</text>
+
+ <!-- Hop 6: Send Alert -> Abort -->
+ <line x1="555" y1="205" x2="615" y2="205" class="wf-hop-failure"
marker-end="url(#arrow-fail)" />
+ <rect x="565" y="190" width="42" height="17" rx="3" fill="#fee2e2" />
+ <text x="586" y="203" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="9.5" font-weight="700"
fill="#dc2626">FALSE</text>
+
+ <!-- ACTION 0: START -->
+ <g transform="translate(25, 105)">
+ <rect width="80" height="90" class="wf-card-start" />
+ <circle cx="40" cy="34" r="14" fill="#0e3a5a" />
+ <polygon points="36,27 48,34 36,41" fill="#ffffff" />
+ <text x="40" y="64" text-anchor="middle" class="wf-text-title"
font-size="13">Start</text>
+ <text x="40" y="78" text-anchor="middle" font-family="system-ui,
-apple-system, sans-serif" font-size="9.5" fill="#64748b">Origin</text>
+ </g>
+
+ <!-- ACTION 1: CHECK SOURCE FILES -->
+ <g transform="translate(165, 105)">
+ <rect width="160" height="90" class="wf-card" />
+ <rect x="12" y="10" width="64" height="18" rx="3" fill="#f1f5f9" />
+ <text x="44" y="23" text-anchor="middle" class="wf-text-tag
wf-ink-muted">Action 1</text>
+ <text x="14" y="50" class="wf-text-title">Check Files</text>
+ <text x="14" y="68" class="wf-text-subtitle">Prerequisites check</text>
+ <text x="14" y="83" font-family="system-ui, -apple-system, sans-serif"
font-size="11" font-weight="700" class="wf-ink-strong">Exit: True / False</text>
+ </g>
+
+ <!-- ACTION 2: RUN INGESTION PIPELINE (SUCCESS BRANCH) -->
+ <g transform="translate(395, 60)">
+ <rect width="160" height="90" class="wf-card-success" />
+ <rect x="12" y="10" width="64" height="18" rx="3" fill="#d1fae5" />
+ <text x="44" y="23" text-anchor="middle" class="wf-text-tag"
fill="#059669">Action 2</text>
+ <text x="14" y="50" class="wf-text-title">Run Pipeline</text>
+ <text x="14" y="68" class="wf-text-subtitle">Fires data worker</text>
+ <text x="14" y="83" font-family="system-ui, -apple-system, sans-serif"
font-size="11" font-weight="700" fill="#059669">Streams records</text>
+ </g>
+
+ <!-- ACTION 3: ARCHIVE FILES (SUCCESS BRANCH) -->
+ <g transform="translate(615, 60)">
+ <rect width="150" height="90" class="wf-card" />
+ <rect x="12" y="10" width="64" height="18" rx="3" fill="#f1f5f9" />
+ <text x="44" y="23" text-anchor="middle" class="wf-text-tag
wf-ink-muted">Action 3</text>
+ <text x="14" y="50" class="wf-text-title">Archive Files</text>
+ <text x="14" y="68" class="wf-text-subtitle">Moves input to S3</text>
+ <text x="14" y="83" font-family="system-ui, -apple-system, sans-serif"
font-size="11" font-weight="700" fill="#64748b">Cleans temp dir</text>
+ </g>
+
+ <!-- SUCCESS END ACTION -->
+ <g transform="translate(805, 85)">
+ <circle cx="20" cy="20" r="18" fill="#10b981" />
+ <polyline points="13,20 18,25 28,15" fill="none" stroke="#ffffff"
stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" />
+ <text x="20" y="50" text-anchor="middle" class="wf-text-title"
font-size="12" fill="#059669">Success</text>
+ </g>
+
+ <!-- ACTION 4: SEND ALERT EMAIL (FAILURE BRANCH) -->
+ <g transform="translate(395, 160)">
+ <rect width="160" height="90" class="wf-card-fail" />
+ <rect x="12" y="10" width="66" height="18" rx="3" fill="#fee2e2" />
+ <text x="45" y="23" text-anchor="middle" class="wf-text-tag"
fill="#dc2626">Fallback</text>
+ <text x="14" y="50" class="wf-text-title">Send Alert</text>
+ <text x="14" y="68" class="wf-text-subtitle">Emails on-call team</text>
+ <text x="14" y="83" font-family="system-ui, -apple-system, sans-serif"
font-size="11" font-weight="700" fill="#dc2626">Logs error details</text>
+ </g>
+
+ <!-- ACTION 5: ABORT WORKFLOW -->
+ <g transform="translate(615, 160)">
+ <rect width="150" height="90" class="wf-card-fail" />
+ <rect x="12" y="10" width="54" height="18" rx="3" fill="#fee2e2" />
+ <text x="39" y="23" text-anchor="middle" class="wf-text-tag"
fill="#dc2626">Abort</text>
+ <text x="14" y="50" class="wf-text-title">Abort Workflow</text>
+ <text x="14" y="68" class="wf-text-subtitle">Flags job as failed</text>
+ <text x="14" y="83" font-family="system-ui, -apple-system, sans-serif"
font-size="11" font-weight="700" fill="#dc2626">Stops execution</text>
+ </g>
+
+ <!-- Sequential Step Execution Tracker (Pulsing Token) -->
+ <circle cx="0" cy="0" r="6" fill="#00a3a6" stroke="#ffffff" stroke-width="2"
class="wf-token-seq" />
+
+ <!-- Bottom Explanatory Ribbon -->
+ <g transform="translate(25, 260)">
+ <rect width="830" height="36" class="wf-ribbon" />
+ <circle cx="20" cy="18" r="4.5" fill="#10b981" />
+ <text x="34" y="22" class="wf-ribbon-text">
+ Backtracking: Hop follows one branch to its end, then returns to the
last fork and takes any branch it has not run yet.
+ </text>
+ </g>
+</svg>
diff --git a/docs/hop-user-manual/modules/ROOT/pages/concepts.adoc
b/docs/hop-user-manual/modules/ROOT/pages/concepts.adoc
index 095be56f33..66d0c9e88f 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/concepts.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/concepts.adoc
@@ -16,10 +16,14 @@ under the License.
////
[[Concepts]]
:imagesdir: ../assets/images
-:description: Hop comes with a number of concepts: a variety of tools, a large
number of metadata types, projects and enviromments. At the core of literally
everything in Hop is metadata.
+:description: Hop comes with a number of concepts: a variety of tools, a large
number of metadata types, projects and environments. At the core of literally
everything in Hop is metadata.
= Concepts
+Apache Hop is designed around a clean separation of concerns: **Workflows**
orchestrate jobs and processes step by step, while **Pipelines** perform
high-throughput streaming data processing in parallel.
+
+image::concepts/hop-architecture-overview.svg[Apache Hop Architecture
Overview,opts=inline]
+
include::snippets/hop-tools/hop-tools.adoc[]
include::snippets/hop-concepts/item-types.adoc[]
@@ -37,9 +41,16 @@ include::snippets/hop-concepts/data-types.adoc[]
xref:pipeline/formatting-values.adoc[Formatting numbers and dates] describes
Format, Length, Precision, Decimal, Group and Currency — the field metadata
used whenever Hop converts a value to or from a String.
+== Core Guides
+
+For in-depth explanations of how Hop executes and orchestrates data solutions:
+
+* xref:pipeline/pipelines.adoc[Pipelines] — Streaming parallel data workers,
in-memory row buffers, automatic back-pressure, execution engines (Local,
Remote, Beam, Spark), and real-world scenarios.
+* xref:workflow/workflows.adoc[Workflows] — Process orchestration, sequential
execution, conditional decision hops, and backtracking path exploration.
+
== Various
-The following items are an alphabetically ordered list of concepts that are
used throughout Hop and will be mentioned at various locations in the Hop tools
and documentation.
+The following items are concepts that are used throughout Hop and will be
mentioned at various locations in the Hop tools and documentation.
Fields, parameters, and variables::
include::snippets/hop-concepts/fields-parameters-variables.adoc[]
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-concepts.adoc
b/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-concepts.adoc
index 7628429e17..983a3d667f 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-concepts.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-concepts.adoc
@@ -17,39 +17,60 @@ under the License.
[[HopConcepts]]
:imagesdir: ../../assets/images
:page-pagination:
-:description: Getting Started (3/9): before going into the details of
workflows and pipelines, we'll walk through the core terminology used in Hop.
+:description: Getting Started (3/9): before going into the details of
workflows and pipelines, we'll walk through the core terminology and
architecture used in Hop.
= Hop Concepts
== Core Concepts
-Before we dive deeper, let's take a minute to familiarize ourselves with the
Hop lingo.
-
-**Metadata** is by far the most important concept in all of Hop.
-Every item we'll cover below is defined as metadata.
-All interactions between Hop and other components in your data architecture
are done through metadata.
-_Metadata is at the core of **everything** in Hop_.
+Before we dive into building, let's take a minute to understand the mental
model behind Apache Hop.
+**Metadata** is at the core of literally everything in Hop.
+Pipelines, workflows, relational database connections, run configurations,
unit tests, and servers are all defined as metadata objects.
+You define a metadata item once (for example, a database connection called
`CRM`) and reference it everywhere by name.
+When a credential or host changes, you update it once in your project metadata
and every pipeline and workflow automatically adopts the change.
+See the xref:metadata-types/index.adoc[full list of metadata types].
-* **Pipelines** are collections of **transforms**, connected by **hops**.
-All transforms in a pipeline run in parallel.
+== Pipelines vs. Workflows: At a Glance
-* **Workflows** are collections of **actions**, connected by **hops**.
-All actions in a workflow run sequentially by default.
+The most critical architectural concept in Hop is the distinction between
**Pipelines** and **Workflows**.
+They may look visually similar in the GUI canvas with icons connected by
arrows, but they operate completely differently:
-* **Metadata items** are the definitions you create once and reuse everywhere:
relational database connections,
xref:pipeline/pipeline-run-configurations/pipeline-run-configurations.adoc[run
configurations], xref:metadata-types/hop-server.adoc[Hop Servers],
xref:metadata-types/data-set.adoc[data sets],
xref:metadata-types/pipeline-unit-test.adoc[unit tests], and many more.
-A pipeline refers to a database connection by name; the connection itself is
defined once, in the project.
-Change it in one place and everything that uses it follows.
-See the xref:metadata-types/index.adoc[full list of metadata types].
+[options="header", cols="25%,37%,38%"]
+|===
+|Concept |Pipeline (`.hpl`) |Workflow (`.hwf`)
+|**Primary Role** |**Data Worker** (Processes & transforms rows)
|**Orchestrator** (Coordinates jobs & tasks)
+|**Building Blocks** |xref:pipeline/transforms.adoc[Transforms]
|xref:workflow/actions.adoc[Actions]
+|**Execution Mode** |**Parallel** (All transforms run simultaneously)
|**Sequential** (Actions run one after another)
+|**What Flows on Hops** |Real **data rows** streaming through in-memory
buffers |**Control flow / Exit status** (Success or Failure)
+|**Starting Points** |One or more xref:pipeline/pipeline-sources.adoc[pipeline
sources] |Exactly one xref:workflow/actions/start.adoc[Start] action
+|**Back-Pressure** |Automatic (upstream pauses if downstream is slow) |N/A
(state machine progresses step by step)
+|**Typical Operations** |Reading tables, filtering, joining streams, cleaning
rows |Checking prerequisites, running pipelines, alerting
+|===
-* **Projects** are logical collections of hop code and configuration.
-**Environments** contain the environment-specific (e.g. dev, uat, prd)
metadata.
+The diagram below illustrates how Workflows serve as the orchestration control
plane, firing parallel Pipelines as the data processing engine across your
infrastructure:
-* **Fields** are the typed columns on a data row.
-**Parameters** are named inputs on a pipeline or workflow.
-**Variables** are named string values in a scope.
-They look similar in the UI but they do not behave the same way — see
xref:fields-parameters-variables.adoc[Fields, parameters, and variables].
+image::concepts/hop-architecture-overview.svg[Apache Hop Architecture Overview
- Workflows and Pipelines,opts=inline]
include::../snippets/hop-concepts/item-types.adoc[]
include::../snippets/hop-concepts/hop-projects-environments.adoc[]
+
+== Fields, Parameters, and Variables
+
+In Hop, values belong to three distinct categories:
+
+* **Fields**: The named and typed columns on a data row in a pipeline (e.g.
`customer_id` as an Integer, `order_date` as a Date). Fields travel across hops
between transforms.
+* **Parameters**: Named inputs declared on a pipeline or workflow. They allow
you to pass external values into a job at runtime (for example, passing a
specific date or processing batch ID).
+* **Variables**: Named environment and system variables (e.g.
`+${PROJECT_HOME}+` or `+${JAVA_HOME}+`). Variables are resolved in text fields
and paths across Hop.
+
+See xref:fields-parameters-variables.adoc[Fields, parameters, and variables]
for a comprehensive guide on their scopes and behaviors.
+
+== Where to Go Next
+
+Now that you know the vocabulary:
+
+* Continue with the xref:getting-started/hop-gui.adoc[Hop Gui Overview] to
explore the user interface.
+* Or dive deeper into the dedicated architecture guides:
+** xref:pipeline/pipelines.adoc[Pipelines Deep Dive] — Learn about parallel
streaming, in-memory row buffers, execution engines, and real-world scenarios.
+** xref:workflow/workflows.adoc[Workflows Deep Dive] — Learn about
orchestration, conditional decision hops, and backtracking path exploration.
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-pipelines.adoc
b/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-pipelines.adoc
index 47b07c60d5..0ba0b13f0e 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-pipelines.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-pipelines.adoc
@@ -21,14 +21,23 @@ under the License.
= Your first pipeline
-In xref:getting-started/hop-concepts.adoc[Concepts] we said a pipeline is a
set of transforms connected by hops, and that all transforms run in parallel.
+== Why and When to Use a Pipeline
-This page turns that into a working pipeline.
-It is deliberately small, but it is a complete one: it creates data, changes
it, lets you look at it, and writes it out.
-That read - change - write shape is what almost every pipeline you build later
will look like.
+In Apache Hop, a **pipeline** is your streaming data worker.
+Whenever your task involves moving, filtering, cleaning, joining, or
transforming actual rows of data, you build a pipeline.
+Whether you are extracting millions of records from a database into a data
lake, listening to real-time events on an Apache Kafka topic, or cleaning
customer addresses from a CSV file, pipelines perform the heavy lifting.
-You will need nothing but a Hop installation.
-No database, no sample files, no network.
+Key things to remember before you build:
+
+* **Transforms are the workers**: Each transform on the canvas does one
specific job (e.g. generating data, calculating a value, writing to disk).
+* **Parallel streaming**: All transforms in a pipeline start at the same time
and run concurrently. Rows stream continuously through in-memory row buffers
between transforms.
+* **Automatic back-pressure**: If a downstream transform is slower than an
upstream transform, the buffer between them fills up and upstream automatically
pauses. This keeps the memory held between transforms bounded, however large
the data set.
+
+TIP: For an architectural deep dive into how pipelines run, row buffers, and
execution engines (Local, Remote, Beam, Spark), see the
xref:pipeline/pipelines.adoc[Pipelines guide].
+
+Now, let's turn these concepts into a working pipeline on your workstation.
+The example below is deliberately self-contained: it creates data in memory,
modifies it, lets you inspect it live with preview, and writes it to a file.
+You will need nothing but your Hop installation — no external databases or
network connections required.
== Create a pipeline
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-workflows.adoc
b/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-workflows.adoc
index 157f861c20..8eff779f1b 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-workflows.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/getting-started/hop-gui-workflows.adoc
@@ -21,17 +21,25 @@ under the License.
= Your first workflow
-A pipeline moves rows.
-A workflow decides *what runs, in what order, and what happens when something
fails*.
+== Why and When to Use a Workflow
-That difference is the point of this chapter.
-We'll build a workflow that runs the pipeline from the
xref:getting-started/hop-gui-pipelines.adoc[previous chapter] and reports a
different message depending on the outcome.
+Where pipelines do the heavy lifting of moving and transforming individual
data rows, **workflows** manage the big picture.
+A workflow is your process coordinator: it decides *what tasks run, in what
order, and what happens when something succeeds or fails*.
-Remember from xref:getting-started/hop-concepts.adoc[Concepts]:
+In a production data platform, you almost never run a pipeline completely
isolated in a vacuum.
+Before a pipeline processes data, you might need to check if a database is
reachable or whether an incoming file has arrived on an SFTP server.
+After the pipeline finishes, you might want to archive the input files, write
an audit log, or send an email alert if an unexpected error occurred.
+Workflows handle all of these orchestration tasks.
-* A **workflow** is a sequence of **actions** connected by **hops**, with a
start point and one or more end points.
-* Unlike transforms in a pipeline, actions run **one after another**, not in
parallel.
-* A **hop** in a workflow can be conditional: it decides which action runs
next based on whether the previous one succeeded.
+Key things to remember before you build:
+
+* **Actions are the steps**: Each action performs an atomic job (e.g. running
a pipeline, checking a file, sending an email).
+* **Sequential execution**: Unlike transforms in a pipeline, actions run **one
after another in sequence** by default.
+* **Conditional decision routing**: Hops in a workflow are conditional
gateways. You can follow different branches depending on whether an action
succeeded (green hop) or failed (red hop).
+
+TIP: For an architectural deep dive into how workflows execute, the
backtracking algorithm, and execution engines (Local, Remote, Load-Balancing),
see the xref:workflow/workflows.adoc[Workflows guide].
+
+In this chapter, we will build a workflow that orchestrates the pipeline you
created in the xref:getting-started/hop-gui-pipelines.adoc[previous chapter]
and takes different paths depending on whether the pipeline succeeded or failed.
== Create a workflow
diff --git a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
index b5d196285f..525d66b592 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
@@ -16,47 +16,166 @@ under the License.
////
[[Pipelines]]
:imagesdir: ../assets/images
-:description: Pipelines, together with workflows, are the main building blocks
in Hop. Pipelines perform the heavy data lifting: in a pipeline, you read data
from one or more sources, perform a number of operations (joins, lookups,
filters and lots more) and finally write the processed data to one or more
target platforms.
+:description: Pipelines perform the heavy data lifting in Apache Hop: reading,
transforming, enriching, cleaning, validating, and writing data across diverse
platforms in a streaming, parallel execution model.
= Pipelines
-== Pipelines overview
+== What Pipelines Actually Do
-Pipelines, together with workflows, are the main building blocks in Hop.
Pipelines perform the heavy data lifting: in a pipeline, you read data from one
or more sources, perform a number of operations (joins, lookups, filters and
lots more) and finally write the processed data to one or more target platforms.
+Pipelines are the **data workers** of Apache Hop.
+Where xref:workflow/workflows.adoc[workflows] orchestrate tasks and make
decisions, pipelines do the heavy data processing: they read data from one or
more sources, transform, enrich, and validate it in-flight, and write the
processed records to one or more destination systems.
-Pipelines are a network of xref:pipeline/transforms.adoc[transforms],
connected by hops. Just like the xref:workflow/actions.adoc[actions] in a
workflow, each transform is a small piece of functionality. The combination of
a number of transforms allow Hop developers to build powerful data processing
and, in combination with workflows, orchestration solutions.
+Pipelines are designed for high throughput and low latency:
-Even though there is some visual resemblance, workflows and pipelines operate
very differently.
+* **Ingestion**: Read structured and unstructured data from relational
databases, cloud object storage (S3, GCS, Azure Blob), local or network files
via the xref:vfs.adoc[Virtual File System (VFS)], message queues (Apache Kafka,
JMS, AWS SQS), or REST APIs.
+* **Transformation & Cleansing**: Filter invalid rows, calculate values,
standardize date and number formats, manipulate strings, parse JSON or XML,
handle missing values, and execute business rules.
+* **Lookups & Joins**: Enrich streaming records by joining data streams or
performing real-time lookups against database tables, in-memory caches, or
static datasets.
+* **Loading**: Insert, update, or bulk load rows into relational databases,
cloud data warehouses (Snowflake, BigQuery, Redshift), lakehouse tables (Delta
Lake, Apache Iceberg), search indexes, or messaging topics.
-The core principles of pipelines are:
+== Under the Hood: How a Pipeline Executes
-* pipelines are networks. Each transform in a pipeline is part of the network.
-* a pipeline runs all of its transforms in parallel. All transforms are
started and process data simultaneously. In a simple pipeline where you read
data from a large file, do some processing and finally write to a database
table, you're typically still reading from the file while you're already
loading data to the database.
-* data flows through the various transforms in a pipeline over hops. In
contrast to workflow hops, pipeline hops typically don't have an exit status.
Pipelines do have some routing capabilities through e.g.
xref:pipeline/transforms/filterrows.adoc[Filter Rows] transform and
xref:pipeline/errorhandling.adoc[error handling], but the core pipeline
principle still applies: the pipeline is a network, and data flow through the
network in parallel.
+Even though a pipeline diagram looks like a flowchart, it does **not** execute
step-by-step.
+A pipeline is an active streaming dataflow network.
-== Example pipeline walk-through
+image::concepts/pipeline-parallel-flow.svg[Pipeline Parallel Streaming and
In-Memory Row Buffers,opts=inline]
-The example below shows a very basic pipeline. This is what happens when we
run this pipeline:
+=== 1. The Parallel Nature
-* the pipeline has 7 transforms. All 7 of these transforms become active when
we start the pipeline.
-* the "read-25M-records" transform starts reading data from a file, and pushes
that data down the stream to "perform-calculations" and the following
transforms. Since reading 25 million records takes a while, some data may
already have finished processing while we're still reading records from the
file.
-* the "lookup-sql-data" matches data we read from the file with data we
retrieved from the "read-sql-data" transform. The
xref:pipeline/transforms/streamlookup.adoc[Stream Lookup] accepts input from
the "read-sql-data", which is shown with the information icon
image:icons/info.svg[] on the hop.
-* once the data from the file and sql query are matched, we check a condition
with the xref:pipeline/transforms/filterrows.adoc[Filter Rows] transform in
"condition?". The output of this data is passed to "write-to-table" or
"write-to-file", depending on whether the condition outcome was true or false.
+When you run a pipeline, **all transforms start at the exact same moment**.
+Each transform operates as an independent worker running simultaneously
alongside every other transform.
-image:hop-gui/pipeline/basic-pipeline.png[Pipelines - basic pipeline,
width="65%"]
+In a traditional batch script, step 1 must finish reading an entire
million-row file before step 2 can start processing.
+In a Hop pipeline:
-== Next steps
+* The input transform begins reading the first rows from your file.
+* It immediately pushes those rows down the hop to the next transform.
+* The downstream transforms begin filtering and calculating right away.
+* While the input transform is reading row 50,000, intermediate transforms are
processing row 25,000, and the output transform is already writing row 10,000
to the database!
-Pipelines are an extensive topic. Check the pages below to learn more about
working with pipelines:
+=== 2. In-Memory Row Buffers
-* xref:pipeline/hop-pipeline-editor.adoc[Pipeline Editor]
-* xref:pipeline/create-pipeline.adoc[Create a Pipeline]
-* xref:pipeline/run-preview-debug-pipeline.adoc[Run, Preview and Debug a
Pipeline]
-*
xref:pipeline/pipeline-run-configurations/pipeline-run-configurations.adoc[Pipeline
Run Configurations]
-* xref:pipeline/metadata-injection.adoc[Metadata Injection]
-* xref:pipeline/partitioning.adoc[Partitioning]
-* xref:pipeline/beam/getting-started-with-beam.adoc[Getting started with
Apache Beam]
-* xref:pipeline/spark/getting-started-with-native-spark.adoc[Getting started
with the native Spark pipeline engine]
-* xref:pipeline/spark/lakehouse.adoc[Lakehouse tables on the native Spark
engine (Delta / Iceberg)]
-* xref:pipeline/pipeline-sources.adoc[Pipeline sources]
-* xref:pipeline/transforms.adoc[Transforms]
+Data travels between transforms along **hops**.
+Behind every hop sits an **in-memory row buffer** (a queue with a default
capacity of 10,000 rows, configurable in the pipeline properties dialog).
+
+* An upstream transform pushes rows into the buffer.
+* The downstream transform pulls rows out of the buffer to do its work.
+* Data streams directly through memory from transform to transform, avoiding
slow intermediate temporary files on disk.
+
+=== 3. Automatic Back-Pressure
+
+What happens when an output database or web service is slower than a fast file
reader?
+Hop handles this automatically with **back-pressure**:
+
+1. If the downstream transform cannot keep up, the row buffer between the two
transforms fills up to its maximum capacity (e.g. 10,000 rows).
+2. Once the buffer is full, the upstream transform automatically pauses
reading until the downstream transform consumes rows and makes room.
+3. This pause naturally ripples backward through the entire pipeline all the
way to the initial source transform.
+4. When the downstream transform catches up, the upstream transform resumes
processing automatically.
+
+Thanks to back-pressure, your pipelines automatically adapt to varying speeds
and the memory held between transforms stays bounded, no matter how many
millions or billions of rows you process.
+Note that back-pressure governs the row buffers on the hops only.
+Transforms that have to collect a whole set before they can emit a row --
xref:pipeline/transforms/streamlookup.adoc[Stream Lookup],
xref:pipeline/transforms/memgroupby.adoc[Memory Group By] and similar -- hold
that set in memory and still need to fit in the heap.
+
+=== 4. Scaling and Partitioning
+
+When a single transform becomes a performance bottleneck (such as a complex
calculation or slow API call), Hop provides built-in tools to scale:
+
+* **Multiple Copies**: You can tell Hop to run multiple parallel copies of any
transform (e.g., 4 or 8 copies). Hop will automatically distribute rows across
these copies to saturate available CPU cores
(`xref:pipeline/specify-copies.adoc[Specify copies]`).
+* **Data Partitioning**: Partition data streams across distinct worker threads
or cluster nodes according to partition rules
(`xref:pipeline/partitioning.adoc[Partitioning]`).
+* **Row-Level Error Handling**: Route rejected or malformed rows into an
alternate error stream without stopping the pipeline
(`xref:pipeline/errorhandling.adoc[Error Handling]`).
+
+== Pipeline Execution Engines
+
+In Apache Hop, your pipeline design (`.hpl`) contains pure metadata and does
not hardcode where or how it runs.
+You choose an engine via a
xref:pipeline/pipeline-run-configurations/pipeline-run-configurations.adoc[Pipeline
Run Configuration].
+
+Hop provides several purpose-built pipeline engines:
+
+[options="header", cols="22%,38%,40%"]
+|===
+|Engine |Best For |How It Operates
+|xref:pipeline/pipeline-run-configurations/native-local-pipeline-engine.adoc[Native
Local]
+|Workstations, local testing, and standard single-server production jobs.
+|Runs locally in the Hop JVM, using multi-threaded in-memory streaming with
row buffers.
+|xref:pipeline/pipeline-run-configurations/native-remote-pipeline-engine.adoc[Native
Remote]
+|Offloading jobs to dedicated servers or cloud instances.
+|Sends pipeline metadata to a remote xref:hop-server/index.adoc[Hop Server]
over HTTP, streaming logs and metrics back to the client.
+|xref:pipeline/pipeline-run-configurations/native-load-balancing-pipeline-engine.adoc[Native
Load-balancing]
+|Spreading separate pipeline runs over a group of servers.
+|Assigns each pipeline to one server from a configured group, using an
even-load or pack algorithm, and retries on another server when the chosen one
is at capacity.
+|xref:pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc[Beam
Flink]
+|Massive batch processing and low-latency continuous streaming at enterprise
scale.
+|Translates the pipeline into an Apache Flink streaming or batch application
running on a Flink cluster.
+|xref:pipeline/pipeline-run-configurations/beam-spark-pipeline-engine.adoc[Beam
Spark]
+|Distributed batch data processing on existing Apache Spark clusters.
+|Compiles the pipeline into an Apache Beam job executed on an Apache Spark
cluster.
+|xref:pipeline/pipeline-run-configurations/beam-dataflow-pipeline-engine.adoc[Beam
Google Dataflow]
+|Serverless, auto-scaling big data pipelines on Google Cloud.
+|Deploys the pipeline to Google Cloud Dataflow, with automatic resource
provisioning and horizontal scaling.
+|xref:pipeline/pipeline-run-configurations/beam-direct-pipeline-engine.adoc[Beam
Direct]
+|Testing Beam pipelines locally before deploying to Flink, Spark, or Dataflow.
+|Executes Apache Beam pipelines locally on your machine for quick verification.
+|xref:pipeline/pipeline-run-configurations/native-spark-pipeline-engine.adoc[Native
Spark]
+|Lakehouse architectures, Delta Lake, and Apache Iceberg tables.
+|Executes natively on Apache Spark using Spark SQL and DataFrames without
requiring Beam.
+|xref:pipeline/pipeline-run-configurations/single-threaded-pipeline-engine.adoc[Single
Threaded]
+|Micro-batches, low-latency API handling, and deterministic debugging.
+|Executes transforms sequentially on a single thread, passing batches of rows
one transform at a time. Ideal for embedded use and sub-pipelines.
+|===
+
+== Popular Real-World Scenarios
+
+Here is how pipelines are commonly used in modern data architectures:
+
+=== 1. High-Throughput Batch Ingestion & Data Warehousing
+Pipelines extract millions of records from operational databases (ERP, CRM) or
files (CSV, Parquet, Excel) stored on S3 or Azure Blob storage.
+The pipeline cleanses fields, checks referential integrity with cached
lookups, and bulk loads data directly into analytical databases like Snowflake,
BigQuery, PostgreSQL, or Redshift.
+
+=== 2. Continuous Real-Time Streaming
+A pipeline listens 24/7 to an Apache Kafka topic using the
xref:pipeline/transforms/kafkaconsumer.adoc[Kafka Consumer] transform.
+As events arrive, the pipeline deserializes JSON or Avro payloads, enriches
them with static metadata, computes rolling metrics, and streams outputs
directly to monitoring dashboards or operational stores.
+
+=== 3. Interactive Workstation Development & Live Preview
+While building data pipelines in the Hop GUI, data engineers do not need to
wait for full batch runs to see if their logic works.
+By right-clicking any transform and selecting **Preview & debug output**, you
can inspect real data rows flowing through that transform immediately.
+See xref:pipeline/run-preview-debug-pipeline.adoc[Run, Preview and Debug a
Pipeline].
+
+=== 4. Automated CI/CD Regression Testing
+Pipelines can be linked with golden datasets using the
xref:pipeline/pipeline-unit-testing.adoc[Pipeline Unit Testing] framework.
+In continuous integration (CI) environments (e.g. GitHub Actions, GitLab CI),
tests run automatically headlessly with `hop-run` to verify that data
transformations produce expected results before code is merged or deployed.
+
+=== 5. Embedded in Custom Applications
+Using Hop's runtime libraries, developers can embed pipelines directly inside
custom Java applications.
+The host application dynamically configures variables, launches the pipeline,
and consumes output data streams without launching the GUI.
+
+=== 6. Python Data Science with PyHop
+Data scientists can use the xref:hop-tools/hop-python/hop-python.adoc[PyHop]
library to control Hop from Python scripts.
+Pipelines can be triggered programmatically, and high-volume data streams can
be passed directly into pandas or Polars DataFrames using Apache Arrow
in-memory IPC without touching disk.
+
+== Example Pipeline Walk-Through
+
+The example below shows a typical real-world pipeline design:
+
+image:hop-gui/pipeline/basic-pipeline.png[Pipelines - basic pipeline,
width="75%"]
+
+When this pipeline runs:
+
+1. All 7 transforms start simultaneously.
+2. The `read-25M-records` transform begins reading from a file and pushes rows
downstream.
+3. The `perform-calculations` transform processes those incoming rows
concurrently.
+4. The `lookup-sql-data` transform enriches the rows with reference data
retrieved from `read-sql-data` using a
xref:pipeline/transforms/streamlookup.adoc[Stream Lookup].
+5. The xref:pipeline/transforms/filterrows.adoc[Filter Rows] transform splits
rows based on a business condition: matching rows go to `write-to-table`, while
non-matching rows go to `write-to-file`.
+6. Rows flow continuously through the buffers until all records are processed.
+
+== Next Steps
+
+To dive deeper into building and managing pipelines, explore:
+
+* xref:pipeline/hop-pipeline-editor.adoc[Pipeline Editor] — Learn your way
around the pipeline visual designer.
+* xref:pipeline/create-pipeline.adoc[Create a Pipeline] — Step-by-step
instructions for creating and saving pipelines.
+* xref:pipeline/pipeline-sources.adoc[Pipeline Sources] — The transforms that
can start a pipeline.
+* xref:pipeline/transforms.adoc[Transforms Catalog] — Browse the complete
library of over 100 transforms.
+* xref:pipeline/run-preview-debug-pipeline.adoc[Run, Preview and Debug a
Pipeline] — Techniques for inspecting live rows in-flight.
+* xref:pipeline/pipeline-metrics.adoc[Pipeline Metrics] — Understand row
counters, speeds, and buffer metrics.
+* xref:pipeline/metadata-injection.adoc[Metadata Injection] — Generate
dynamic, metadata-driven pipelines at runtime.
+* xref:pipeline/beam/getting-started-with-beam.adoc[Getting Started with
Apache Beam] — Scale out pipelines on Flink, Spark, and Google Cloud Dataflow.
+* xref:pipeline/spark/getting-started-with-native-spark.adoc[Getting Started
with Native Spark] — Run natively on Spark clusters and lakehouse tables.
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/action.adoc
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/action.adoc
index 5242b4f20b..54690d8daf 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/action.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/action.adoc
@@ -14,6 +14,9 @@ KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
////
-An Action is one operation performed in a Workflow.
-Actions are executed sequentially by default, with parallel execution as a
configuration option.
-An Action returns a true or false exit code, which can be used (or ignored) in
the Workflow’s execution.
\ No newline at end of file
+An Action is a single operation performed in a Workflow.
+Actions are executed sequentially one after another by default, with parallel
branch execution available as a configuration option.
+When an action finishes, it returns a success or failure status (true or
false).
+This status determines which outgoing hops the workflow follows next.
+Actions can also pass variables, lists of files, and summary rows to
subsequent actions.
+See the xref:workflow/actions.adoc[Actions catalog] for all available workflow
actions.
\ No newline at end of file
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/item-types.adoc
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/item-types.adoc
index 82874985d3..5662ad5d98 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/item-types.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/item-types.adoc
@@ -25,7 +25,7 @@ include::hop.adoc[]
Pipeline::
include::pipeline.adoc[]
-image::concepts/pipeline.png[Pipeline]
+image::concepts/pipeline-parallel-flow.svg[Pipeline streaming execution and
back-pressure,opts=inline]
Transform::
include::transform.adoc[]
@@ -33,6 +33,4 @@ include::transform.adoc[]
Workflow::
include::workflow.adoc[]
-image::concepts/workflow.png[Workflow]
-
-
+image::concepts/workflow-backtracking-flow.svg[Workflow sequential
orchestration and backtracking,opts=inline]
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/pipeline.adoc
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/pipeline.adoc
index 8f2f17f34a..9bccc1a808 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/pipeline.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/pipeline.adoc
@@ -14,7 +14,8 @@ KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
////
-Pipelines are the actual data workers.
-Operations in a Pipeline read, modify, enrich, clean and write data.
-A pipeline starts with one or more
xref:pipeline/pipeline-sources.adoc[pipeline sources] and runs every transform
in parallel.
-Orchestration of Pipelines is done through othere Pipelines and/or Workflows.
\ No newline at end of file
+Pipelines are the **data workers** of Apache Hop.
+Operations in a pipeline read, transform, enrich, clean, validate, and write
data.
+A pipeline starts with one or more
xref:pipeline/pipeline-sources.adoc[pipeline sources] and runs every transform
in **parallel**.
+Data travels between transforms along hops through in-memory row buffers,
ensuring high throughput and automatic back-pressure.
+See xref:pipeline/pipelines.adoc[Pipelines] for detailed execution mechanics,
engines, and production scenarios.
\ No newline at end of file
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/transform.adoc
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/transform.adoc
index 9bdac74fa4..1318066a25 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/transform.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/transform.adoc
@@ -14,8 +14,8 @@ KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
////
-A Transform is a unit of work performed in a Pipeline.
-Typical Transform operations are reading data from files, databases,
performing lookups or joins, enriching, cleaning data and more.
-All transforms in a Pipeline are executed in parallel.
-xref:pipeline/pipeline-sources.adoc[Pipeline sources] produce rows without
incoming hops; most other transforms wait for those rows.
-Transforms process data and move batches of processed data on Hops for
processing by subsequent Actions.
\ No newline at end of file
+A Transform is a single unit of data processing inside a Pipeline.
+Typical transform operations include reading from databases or files,
filtering rows, calculating values, performing lookups, joining datasets, and
loading to destinations.
+All transforms in a pipeline run simultaneously in parallel.
+xref:pipeline/pipeline-sources.adoc[Pipeline sources] generate rows without
needing an incoming hop, while intermediate transforms continuously process
streaming rows and pass them along hops to downstream transforms through
in-memory row buffers.
+See the xref:pipeline/transforms.adoc[Transforms catalog] for all available
transforms.
\ No newline at end of file
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/workflow.adoc
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/workflow.adoc
index ee7b3949a8..c700773a99 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/workflow.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/snippets/hop-concepts/workflow.adoc
@@ -14,7 +14,8 @@ KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
////
-A Workflow is a sequence of operations that are performed sequentially by
default (with optional parallel execution).
-Workflows usually do not operate on the data directly, but perform
orchestration tasks.
-Typical tasks in a Workflow consist of retrieving and archiving data, sending
emails, error handling etc.
-)
\ No newline at end of file
+A Workflow is the **orchestrator** of your data architecture.
+Workflows execute a series of actions sequentially by default (with optional
parallel branches).
+Workflows manage control flow rather than streaming individual rows: they
check prerequisites, prepare environments, execute pipelines, manage files,
handle errors, and send notifications.
+Hops in a workflow decide which path to follow based on whether an action
succeeded or failed.
+See xref:workflow/workflows.adoc[Workflows] for in-depth orchestration
concepts, decision backtracking, and scheduling scenarios.
\ No newline at end of file
diff --git a/docs/hop-user-manual/modules/ROOT/pages/workflow/workflows.adoc
b/docs/hop-user-manual/modules/ROOT/pages/workflow/workflows.adoc
index 9319dec3bc..b83e31bc10 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/workflow/workflows.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/workflow/workflows.adoc
@@ -16,45 +16,151 @@ under the License.
////
[[Workflows]]
:imagesdir: ../assets/images
-:description: Workflows are one of the core building blocks in Apache Hop.
Where pipelines do the heavy data lifting, workflows take care of the
orchestration work: prepare the environment, fetch remote files, perform error
handling and executing child workflows and pipelines.
+:description: Workflows are the orchestrators in Apache Hop: managing
execution order, verifying environments, coordinating child pipelines, handling
errors, and routing tasks sequentially with smart decision backtracking.
= Workflows
-== Workflow overview
+== What Workflows Actually Do
-Workflows are one of the core building blocks in Apache Hop. Where pipelines
do the heavy data lifting, workflows take care of the orchestration work:
prepare the environment, fetch remote files, perform error handling and
executing child workflows and pipelines.
+Workflows are the **orchestrators** of Apache Hop.
+Where xref:pipeline/pipelines.adoc[pipelines] are the streaming workers that
process rows of data, workflows are the managers that coordinate jobs: they
ensure prerequisites are met, fetch required files, trigger child pipelines and
workflows in the proper order, react to errors, and send notifications.
-Workflows consist of a series of xref:workflow/actions.adoc[actions],
connected by hops. Just like a xref:pipeline/transforms.adoc[transform] in a
xref:pipeline/pipelines.adoc[pipeline], each action is a small piece of
functionality. The combination of a number of actions allows Hop developers to
build powerful data orchestration solutions.
+Workflows manage the operational lifecycle of your data platform:
-Even though there is some visual resemblance, workflows and pipelines operate
very differently.
+* **Environment Preparation & Health Checks**: Verify that target databases
are reachable, ping remote servers, check whether webservices are online, and
verify disk space before launching heavy data jobs.
+* **File & Storage Lifecycle**: Download files via SFTP or cloud storage (S3,
GCS, Azure Blob), decompress zip archives, verify file existence, decrypt
sensitive files with PGP, move processed files to archive folders, and purge
temporary data.
+* **Coordinating Child Pipelines & Workflows**: Launch pipelines using the
xref:workflow/actions/pipeline.adoc[Pipeline] action, pass parameters and
variables down to child tasks, and collect execution status and summary rows.
+* **Resilient Error Handling & Recovery**: If a pipeline or action fails, a
workflow can follow an alternate error route: retry the failed action, run a
cleanup script, roll back transactions, and send alerts.
+* **Alerting & Auditing**: Send email notifications, dispatch Slack messages,
post Nagios or SNMP traps, and log audit information to tracking tables.
-* Workflows perform orchestration tasks. Actions in a workflow usually do not
operate on the data directly (even though you _can_ change data e.g. through
xref:workflow/actions/sql.adoc[SQL]).
-* Workflows have one (and only one) mandatory starting point (a
xref:workflow/actions/start.adoc[Start] action), but can have multiple end
actions.
-* Workflows can
-* Workflows work sequentially by default. Each action in a workflow has a
position in the workflow sequence, and needs to wait before the previous
actions have completed before it starts.
-* Workflow actions do not pass data over hops. Each workflow action has a
`success` or `failure` exit status. This exit status is used to choose the
routing through the workflow.
-* Hops between actions in a workflow have a status: depending on the exit
status of the previous action, a workflow hop can follow the success (green),
failure (orange) or unconditional (black) hop. An unconditional hop ignores the
exit status of the previous action and is followed whether the previous action
failed or succeeded.
+== Under the Hood: How a Workflow Executes
-== Example workflow walk-through
+Workflows operate very differently from pipelines.
+While a pipeline is an active network of parallel streaming transforms, a
workflow is a **deterministic state machine** that executes step by step.
-Like all workflows, the example workflow shown below starts with the `start`
action.
+image::concepts/workflow-backtracking-flow.svg[Workflow Sequential Execution
and Backtracking,opts=inline]
-The Start action is just a placeholder that can't really fail, so the hop out
of a start action is unconditional.
+=== 1. The Sequential Nature
-The workflow then continues with a
xref:workflow/actions/pipeline.adoc[pipeline] action, "first-pipeline". As the
name implies, this action executes a pipeline.
+Workflows run **one action at a time in sequence** by default:
-If "first-pipeline" runs successfully, the workflow continues to
"second-pipeline". If "first-pipeline" fails, the failure hop to
"handle-errors" is followed.
+* Execution begins at a single mandatory starting point: the
xref:workflow/actions/start.adoc[Start] action.
+* An action executes its complete task (e.g. checking a file or running a
pipeline) and waits until it finishes.
+* Upon completion, the action generates an **exit status** indicating whether
it succeeded or failed.
+* The workflow engine inspects the exit status and uses it to choose which
outgoing hops to follow next.
-In this hypothetical example, we don't care about the result of "Second
pipeline", and want to continue to "delete-tmp-files", where any temporary
files are removed.
+=== 2. Conditional Hops & Decision Routing
-If the temporary files are removed successfully, we move on to the "success"
action. Similar to the Start action, success is a visual indicator of
successful completion of this part of the workflow. It's not mandatory and
doesn't add any functionality, but it often is a good visual indicator of an
end point of your workflow's main stream.
+Unlike pipeline hops (which continuously stream data rows), hops between
actions in a workflow act as **conditional decision gateways**:
-image:hop-gui/workflow/basic-workflow.png[Workflows - basic workflows,
width="65%"]
+[options="header", cols="20%,20%,60%"]
+|===
+|Hop Type |Visual Indicator |When the Next Action Runs
+|**Unconditional Hop** |Solid Black |**Always**, regardless of whether the
previous action succeeded or failed.
+|**Success Hop** |Solid Green |**Only when the previous action succeeded**
(returned true / zero errors).
+|**Failure Hop** |Dashed Red / Orange |**Only when the previous action
failed** (returned false or encountered an error).
+|===
-== Next steps
+By combining success and failure hops, you can build self-healing data
workflows that automatically recover from network glitches, missing files, or
data quality exceptions.
-The following pages take you deeper into the process of building and running
workflows:
+=== 3. The Backtracking Path Exploration Algorithm
-** xref:workflow/create-workflow.adoc[Create a Workflow]
-** xref:workflow/run-debug-workflow.adoc[Run and Debug a Workflow]
-**
xref:workflow/workflow-run-configurations/workflow-run-configurations.adoc[Workflow
Run Configurations]
+What happens when an action connects to multiple outgoing hops, or when a
workflow branches into different paths?
+
+Hop uses a **depth-first backtracking algorithm** to traverse the workflow
graph:
+
+1. **Branch Evaluation**: When an action completes, Hop evaluates all outgoing
hops that match its exit status.
+2. **Depth-First Traversal**: Hop selects the first valid path and follows it
sequentially to its final terminus.
+3. **Backtracking**: When that branch reaches an end point (or encounters a
blocked condition), the workflow engine **backtracks** to the last decision
fork and begins traversing any other valid branches that have not yet run.
+4. **Result Aggregation**: Once all valid branches have finished, Hop combines
the results, error counts, and logs from every explored branch into the overall
workflow result.
+
+=== 4. Variables, Parameters, and Data Passing
+
+While workflows do not stream individual rows over hops, they provide powerful
tools for passing context:
+
+* **Setting Dynamic Variables**: Use the
xref:workflow/actions/setvariables.adoc[Set Variables] action to read
configuration files, extract runtime parameters, and make variables available
to all downstream child pipelines and workflows.
+* **Passing Parameters**: Pass parameters down to child pipelines and child
workflows explicitly, keeping your jobs modular and reusable
(`xref:variables/parameter-passing.adoc[Passing parameters to a child]`).
+* **Result Rows & Files**: When a pipeline produces a small summary list of
rows (via `Copy rows to result`), a workflow can capture those rows and loop
over them using child workflows
(`xref:how-to-guides/loops-in-apache-hop.adoc[Loops in Apache Hop]`).
+* **Parallel Action Branches**: When an action must trigger multiple
independent tasks simultaneously, you can enable parallel execution on that
action (`xref:how-to-guides/workflows-parallel-execution.adoc[Parallel
execution in workflows]`), and optionally re-synchronize them using the
xref:workflow/actions/join.adoc[Join action].
+
+== Workflow Execution Engines
+
+Like pipelines, workflows are pure metadata (`.hwf`) and are decoupled from
where they execute through a
xref:workflow/workflow-run-configurations/workflow-run-configurations.adoc[Workflow
Run Configuration].
+
+Hop provides three workflow execution engines:
+
+[options="header", cols="25%,35%,40%"]
+|===
+|Engine |Best For |How It Operates
+|xref:workflow/workflow-run-configurations/native-local-workflow-engine.adoc[Native
Local]
+|Workstations, local testing, and standard production servers.
+|Executes actions sequentially on the local computer or server within the
local Hop JVM.
+|xref:workflow/workflow-run-configurations/native-remote-workflow-engine.adoc[Native
Remote]
+|Offloading orchestration to dedicated servers.
+|Dispatches the workflow to execute on a remote xref:hop-server/index.adoc[Hop
Server] via HTTP, streaming execution progress and logs back to the caller.
+|xref:workflow/workflow-run-configurations/native-load-balancing-workflow-engine.adoc[Native
Load-balancing]
+|Spreading separate workflow runs over a group of servers.
+|Assigns each workflow to one server from a configured group, using an
even-load or pack algorithm, and retries on another server when the chosen one
is at capacity.
+|===
+
+== Popular Real-World Scenarios
+
+Here is how workflows are commonly applied across enterprise data platforms:
+
+=== 1. Production Nightly Batch Orchestration
+A master workflow executes every night to run your organization's core ELT
batch:
+
+1. Connects to database servers and validates connectivity (`Check Db
connections`).
+2. Checks whether new source data files have landed in cloud storage (`Checks
if files exists`).
+3. Executes a series of ingestion pipelines in order, passing the current
processing batch date as a parameter.
+4. If ingestion succeeds, runs the dimensional modeling or Data Vault
pipelines.
+5. Archives processed input files to long-term S3 / Azure cold storage (`Move
files`).
+6. If any step fails, follows the failure hop to notify the on-call data
engineering team via email or Slack (`Mail`).
+
+=== 2. Containerized Workflows on Kubernetes & Docker
+Workflows are frequently packaged into Docker containers and scheduled as
Kubernetes CronJobs or cloud container tasks.
+The container starts headlessly using the `hop-run` CLI command (`hop-run -j
MyProject -e production -f /path/to/workflow.hwf`), executes the workflow, logs
output to stdout, and exits with a standard status code (0 for success,
non-zero for failure).
+See xref:docker-container.adoc[Hop in Docker] and
xref:hop-run/index.adoc[hop-run].
+
+=== 3. External Orchestration with Apache Airflow
+Many organizations use Apache Airflow as an enterprise scheduler while using
Apache Hop for data processing.
+Airflow DAGs trigger Hop workflows using the BashOperator, DockerOperator, or
KubernetesPodOperator.
+The Hop workflow manages its internal conditional branching and pipeline
execution, reporting overall status back to Airflow.
+See xref:how-to-guides/run-hop-in-apache-airflow.adoc[Run Hop in Apache
Airflow].
+
+=== 4. Automated Regression Testing in CI/CD
+Workflows can automate your project's testing strategy before any code is
promoted to production.
+Using the xref:workflow/actions/runpipelinetests.adoc[Run pipeline unit tests]
action, a workflow discovers and executes every unit test in your Hop project.
+If any pipeline transformation produces unexpected test results, the action
fails, halting the CI/CD pipeline and preventing bad code from being deployed.
+
+=== 5. Python-Driven Workflow Execution with PyHop
+Data applications written in Python can load, configure, and execute Hop
workflows programmatically using
xref:hop-tools/hop-python/hop-python.adoc[PyHop].
+The Python application can pass dynamic runtime parameters, start workflow
execution, monitor real-time progress, and evaluate the final result.
+
+== Example Workflow Walk-Through
+
+The example below shows a typical real-world workflow design:
+
+image:hop-gui/workflow/basic-workflow.png[Workflows - basic workflow,
width="75%"]
+
+When this workflow executes:
+
+1. Execution starts at the **Start** action, which can never fail, and follows
the unconditional hop.
+2. The `first-pipeline` action runs its designated pipeline.
+3. If `first-pipeline` succeeds, the workflow follows the green success hop to
`second-pipeline`.
+4. If `first-pipeline` fails, the workflow follows the orange/red failure hop
to `handle-errors`, skipping `second-pipeline`.
+5. After `second-pipeline` runs, the workflow moves along an unconditional hop
to `delete-tmp-files` to ensure temporary files are removed regardless of
outcome.
+6. If cleanup succeeds, execution arrives at the `success` action, marking the
entire workflow as successfully completed.
+
+== Next Steps
+
+To dive deeper into building and managing workflows, explore:
+
+* xref:workflow/create-workflow.adoc[Create a Workflow] — Step-by-step
instructions for creating, configuring, and saving workflows.
+* xref:workflow/run-debug-workflow.adoc[Run and Debug a Workflow] — How to
execute workflows in Hop GUI, view logs, and inspect action metrics.
+* xref:workflow/actions.adoc[Workflow Actions Catalog] — Browse the complete
catalog of over 80 workflow actions.
+*
xref:workflow/workflow-run-configurations/workflow-run-configurations.adoc[Workflow
Run Configurations] — Configure local, remote, and clustered workflow
execution.
+* xref:how-to-guides/scheduling-workflows-and-pipelines.adoc[Scheduling
Workflows and Pipelines] — Built-in repeating schedules and external scheduling
options.
+* xref:how-to-guides/workflows-parallel-execution.adoc[Parallel Execution in
Workflows] — How and when to fork parallel action branches.
+* xref:how-to-guides/loops-in-apache-hop.adoc[Loops in Apache Hop] —
Techniques for looping over files, records, or dates in workflows.
+* xref:how-to-guides/avoiding-deadlocks.adoc[Avoiding Deadlocks] — Best
practices for designing reliable, deadlock-free orchestration graphs.